Skip to content
NBDC Human Database

No datasets in the cart.

Due to system maintenance, the application system, application review by the Data Access Committee will be unavailable during the following period.
Schedule: October 5th (Mon), 2026, 9:00 - October 7th (Wed), 2026, 15:00 (JST)
We apologize for any inconvenience this may cause and appreciate your understanding.

We are currently receiving a large number of applications for data submission, and the review process is taking longer than usual.We sincerely apologize for the delay and kindly ask for your understanding. When submitting an application, we would greatly appreciate it if you could allow sufficient time for the processing.

Following a change to our organizational structure effective April 1, 2026, this division has been renamed from the "Database Center for Life Science, Joint Support-Center for Data Science Research" to the "Database Division for Life Science (DBCLS), BioData Science Initiative (BSI), National Institute of Genetics (NIG)". Where the former name still appears in the guidelines, please read it as the new name.

Dataset ID

NHA000210

Type of data
GWAS for severe COVID-19
Access criteria
Unrestricted-access
Total data volume
272 MB
File formats
  • TXT
  • ZIP
Research
hum0343
Date published
2026-08-13
Date modified
2026-08-13
Secondary ID
hum0343.v5.covid19.v1

Unrestricted-access files linked to this dataset

FileLabelSizeCopy URL
hum0343.v5.covid19.v1.zip272 MB
hum0343.v5.covid19.v1_dictionary_file.txtDictionary file421 B

Analysis method

Genome wide SNPs

Materials and participants
severe COVID-19 (ICD-10: U071): 3,087 cases
Healthy controls: 55,896 individuals
  • Health status
    Mixed
  • Subject count
    58,983 (Individual)
Disease
severe COVID-19 (U071)
Sample description
DNAs extracted from peripheral blood cells
  • Tissue
    Peripheral blood
Sample provider
N/A
Experimental method
Genotyping by array
Target
N/A
Reagent kit
Infinium Asian Screening Array Kit
Platform
Illumina Infinium Asian Screening Array
QC and filtering
Sample QC: We excluded samples with
(1) sample call rate < 0.98
(2) deviation from the East Asian cluster based on PCA
(3) duplicate/twin samples
Variant QC: We excluded variants with
(1) variant call rate < 0.99
(2) allele count < 5
(3) Hardy-Weinberg equilibrium P < 1.0 × 10^-6
(4) >5% allele frequency deviation from the Japanese reference population
Genotype imputation was performed using an in-house Japanese whole-genome sequencing reference panel (n = 11,754). Variants with imputation INFO > 0.7 and allele frequency > 0.005 were retained. Genome-wide association analysis was conducted using REGENIE with age, sex, age × age, age × sex, and PC1-10 as covariates.
Imputation
haplotype phasing: SHAPEIT4
imputation: Minimac4
Analysis method
genotyping: GenomeStudio
association analysis: REGENIE
Variant count
8,993,355
Processed data type
Imputed genotype data
Phenotype data
Included