Due to system maintenance, the application system, application review by the Data Access Committee will be unavailable during the following period.
Schedule: October 5th (Mon), 2026, 9:00 - October 7th (Wed), 2026, 15:00 (JST)
We apologize for any inconvenience this may cause and appreciate your understanding.
We are currently receiving a large number of applications for data submission, and the review process is taking longer than usual.We sincerely apologize for the delay and kindly ask for your understanding. When submitting an application, we would greatly appreciate it if you could allow sufficient time for the processing.
Following a change to our organizational structure effective April 1, 2026, this division has been renamed from the "Database Center for Life Science, Joint Support-Center for Data Science Research" to the "Database Division for Life Science (DBCLS), BioData Science Initiative (BSI), National Institute of Genetics (NIG)". Where the former name still appears in the guidelines, please read it as the new name.
Dataset ID
NHA000190
Type of data
GWAS for gut microbiome GWAS for plasma metabolite GWAS for KEGG Gene Ortholog and KEGG Pathway
524 Japanese individuals (423 species in the gut microbiome) 306 Japanese individuals (306 plasma metabolites) 524 Japanese individuals (KEGG Gene Ortholog and KEGG Pathway)
Subject count
524 (Individual)
Population
Japanese
Sample description
DNAs extracted from peripheral blood cells
Tissue
Peripheral blood
Tumor / normal
Normal
Experimental method
Genotyping by array WGS
Reagent kit
Infinium Asian Screening Array Kit KAPA Hyper Prep Kit TruSeq DNA PCR-Free Library Prep Kit
Platform
Illumina HiSeq 2500 Illumina HiSeq 3000 Illumina HiSeq X Illumina Infinium Asian Screening Array Illumina NovaSeq 6000
Reference genome
GRCh37
QC and filtering
SNP array data: Sample QC: We excluded individuals with low genotyping call rates (call rate < 98%). We included individuals of the estimated Asian ancestry using PCA. Variant QC: We excluded variants with (1) genotyping call rate < 99%, (2) minor allele count < 5, (3) P-value for Hardy-Weinberg equilibrium < 1.0 × 10^−10, and (4) > 5% allele frequency difference compared with the imputation reference panel or the allele frequency panel of Tohoku Medical Megabank Project. Post-imputation QC: We excluded imputed variants with Rsq < 0.7 and minor allele frequency < 1%. WGS: We excluded variants with genotype call rate <90%, ExcessHet > 60, Hardy-Weinberg P<1.0×10−10 After imputation with Beagle v5.1, we excluded imputed variants with minor allele frequency < 1%.