Skip to content
NBDC Human Database

No datasets in the cart.

Due to system maintenance, the application system, application review by the Data Access Committee will be unavailable during the following period.
Schedule: October 5th (Mon), 2026, 9:00 - October 7th (Wed), 2026, 15:00 (JST)
We apologize for any inconvenience this may cause and appreciate your understanding.

We are currently receiving a large number of applications for data submission, and the review process is taking longer than usual.We sincerely apologize for the delay and kindly ask for your understanding. When submitting an application, we would greatly appreciate it if you could allow sufficient time for the processing.

Following a change to our organizational structure effective April 1, 2026, this division has been renamed from the "Database Center for Life Science, Joint Support-Center for Data Science Research" to the "Database Division for Life Science (DBCLS), BioData Science Initiative (BSI), National Institute of Genetics (NIG)". Where the former name still appears in the guidelines, please read it as the new name.

Dataset ID

NHA000182

Type of data
NGS (WGS)
Access criteria
Unrestricted-access
Total data volume
69.2 GB
File formats
  • VCF
  • Markdown
Research
hum0331
Date published
2023-02-01
Date modified
2023-02-01
Secondary ID
hum0331.v1.freq.v1

Unrestricted-access files linked to this dataset

FileLabelSizeCopy URL
NCBN-freeze2.sampleQC.GTfilter.freq.vcf.gz69.2 GB
README.mdREADME2.4 KB

Analysis method

WGS

Materials and participants
Healthy individuals without any cancers or rare diseases (ICD10: Z006): 9290 individuals
  • Health status
    Healthy
  • Subject count
    9290 (Individual)
Disease
Healthy individuals without any cancers or rare diseases (Z006)
Sample description
DNAs extracted from peripheral blood cells
  • Tissue
    Peripheral blood
  • Tumor / normal
    Normal
Sample provider
N/A
Experimental method
WGS
Target
N/A
Reagent kit
TruSeq DNA PCR-Free Library Prep Kit
Fragmentation
Ultrasonic fragmentation
Platform
Illumina NovaSeq 6000
Read type
Paired-end
Read length
150 bp
Reference genome
GRCh38
Mapping
bwa mem (v0.7.15) compatible algorithm (Parabricks 3.1.0 fq2bam)
Read deduplication
MarkDuplicates (GATK4.1.0) compatible algorithm (Parabricks 3.1.0 fq2bam)
Realignment and base quality recalibration
N/A
Mapping quality
No hard filtering was performaed by mapping quality.
QC and filtering
Whole genome sequencing analysis was performed under the following conditions.
- Confirm library size is 400bp-750bp.
- At least 75% of the bases are QV30 or better.
- Total number of bases after removal of duplicate reads by FASTQC is more than 90 GBase.

After alignment and variant calling, the following samples were excluded from the analysis.
- Samples with abnormal values for depth and mapping rate.
- Samples where the depth of the sex chromosome is inconsistent with the clinical information.
- Any of the samples determined to be within the second degree of kinship in the KING program.

Variant call results were filtered for the following
- Genotypes with GQ64 or with less than 25% minor alleles in heterozygous calls are set to no call
- Set VQSR results to FILTER field in VCF
- Set LowCR in FILTER field for variants with less than 95% call rate
- Variants with a Hardy-Weinberg equilibrium test P-value less than 10-6 have HWE set in FILTER field
Analysis method
HaplotypeCaller (GATK 4.1.0) compatible algorithm (Parabricks 3.1.0 haplotypecaller)
Coverage (depth)
Autosomes: 34x
Variant count
Autosomes: 18,899,392
X chromosome: 836,126
Autosomes: 153,554,029
X chromosome: 6,325,046