Skip to content
NBDC Human Database

No datasets in the cart.

Due to system maintenance, the application system, application review by the Data Access Committee will be unavailable during the following period.
Schedule: October 5th (Mon), 2026, 9:00 - October 7th (Wed), 2026, 15:00 (JST)
We apologize for any inconvenience this may cause and appreciate your understanding.

We are currently receiving a large number of applications for data submission, and the review process is taking longer than usual.We sincerely apologize for the delay and kindly ask for your understanding. When submitting an application, we would greatly appreciate it if you could allow sufficient time for the processing.

Following a change to our organizational structure effective April 1, 2026, this division has been renamed from the "Database Center for Life Science, Joint Support-Center for Data Science Research" to the "Database Division for Life Science (DBCLS), BioData Science Initiative (BSI), National Institute of Genetics (NIG)". Where the former name still appears in the guidelines, please read it as the new name.

Dataset ID

JGAD000625

Type of data
bam/gvcf data of NGS (WGS)
Access criteria
Controlled-access (Type II)
Total data volume
17 B
File formats
  • TXT
Research
hum0184
Date published
2022-02-17
Date modified
2022-02-17

Analysis method

WGS

Materials and participants
4,566 Japanese general residents
  • Subject count
    4566 (Individual)
  • Population
    Japanese
Sample description
fastq files of JGAD000338
  • Tumor / normal
    Normal
Experimental method
WGS
Platform
Illumina HiSeq 2500
Illumina NovaSeq 6000
Reference genome
GRCh37
Mapping
BWA mem 0.7.12
Read deduplication
Picard 2.10.6
Realignment and base quality recalibration
GATK 3.7
Mapping quality
Reads with MAPQ< 20 were excluded at variant calling with GATK 3.7 HaplotypeCaller
QC and filtering
Data with bad base quality and high %GC content were removed.
Aligment:
Data matched for the following condition were removed.
- Low mapping rate
- Different insert size
- Gender information mismatch between meta-data and genotype data
- Suspected sex chromosome aberration
Genotyping:
GATK's best practices includes a variant filtering step following Variant Quality Score Recalibration (VQSR)
- DP/GP (DP < 5, GQ < 20, DP > 60, GQ < 95)
- Heterozygosity (F>=0.05)
- Hardy-Weinberg equilibrium (p < 10^-6)
- Repeat & Low Complexity
Principal Component Analysis (PCA):
PCA was performed with individuals included in the 1000 genomes project and outliers from Japanese cluster were removed.
After these filtering steps, variants located in the regions listed as the HighConfidenceRegion (Genome-In-A-Bottle project) were flagged.
Analysis method
GATK 3.7 HaplotypeCaller
Coverage (depth)
HiSeq 2500: 31.8x
NovaSeq 6000: 28x
Variant count
Autosomes: 10,202,908
X chromosome: 410,435
Autosomes: 76,768,387
X chromosome: 2,898,518
Data summary
Whole genome sequencing analyzed data included in the JGAD000338 were mapped to the GRCh37 reference genome sequence, and variant detection was carried out using the GATK (Genome Analysis Toolkit) standards. This project is an initiative of the GEnome Medical alliance Japan (GEM Japan, GEM-J). Learn more