Dataset ID
JGAD000660
- Type of data
- Imputation data and index data for 180,882 patients from BBJ 1st cohort
- Access criteria
- Controlled-access (Type I)
- Total data volume
- 11.1 TB
- File formats
- VCF
- TBI
- Research
- hum0311
- Date published
- 2022-08-12
- Date modified
- 2022-08-12
- DDBJ Search
- JGAD000660 (opens in a new tab)
- JGA Study
- JGAS000541 (opens in a new tab)
Analysis method
Genotyping by array
- Materials and participants
- 180,882 patients from BBJ 1st cohort
ICD10: A15-A16, B16-B17.0, B18.0-B18.1, B17.1, B18.2, C15, C16, C18, C22, C23-C24, C25, C33-C34, C50, C53, C54, C56, C61, C81, D25, E05, E10, E78.0-E78.5, G12, G40-G41, H25-H26, H40-H42, I20, I21-I22, I44-I49, I50, I60, I69.0, I63, I69.3, I70, J30, J41-J44, J45-J46, J80-J84, K05, K74.3-K74.6, L00-L99, L20, M05-M06, M80-M82, N04, N20-N23, N80, R00-R9 - Health statusAffected
- Subject count180,882 (Individual)
- PopulationEast Asian
- Sample description
- DNAs extracted from peripheral blood cells or saliva
- TissuePeripheral blood, Saliva
- Tumor / normalNormal
- Sample provider
- N/A
- Experimental method
- Genotyping by array
- Target
- N/A
- Reagent kit
- HumanExome BeadChip Kit
HumanOmniExpress BeadChip Kit
HumanOmniExpressExome BeadChip Kit - Platform
- Illumina HumanExome
Illumina HumanOmniExpress
Illumina HumanOmniExpressExome - Reference genome
- GRCh38
- QC and filtering
- Before imputation, we excluded SNPs using the following criteria:
Heterozygosity count for each chip < 5
P-value for Hardy-Weinberg equilibrium (HWE) for each chip < 1.0 x 10^-6
*- Genotype concordance rate with whole-genome sequencing (WGS) for 939 samples < 99.5% and its non-reference discordance rate >= 0.5%
Lower call rate SNPs if the position was the same when merging datasets
Call rate < 99%*
P-values for chrX SNPs were calculated by using female samples
We also excluded samples using the following criteria:
Call Rate < 98%
Samples whose inferred sex was not matched with the clinical information
Lower call rate samples for duplicated or monozygotic twin in the dataset
Outliers from East Asian clusters from principal component analysis with 1KGp3v5 samples. - Imputation
- Eagle software (v2.4.1) without a reference panel
Minimac4 software (v1.0.2) - Analysis method
- GenomeStudio Software
- Variant count
- Autosomes: 515,587 SNVs
X chromosome: 11,140 SNVs - Phenotype data
- Included
- Data use policy
- NBDC data sharing policy (JGAP000001)