{"id":"JGAD000220","research":"hum0014","url":"https://humandbs.dbcls.jp/dataset/JGAD000220","datePublished":"2020-09-28","dateModified":"2022-08-23","values":[{"key":"access-criteria","label":{"ja":"アクセス制限","en":"Access criteria"},"type":"vocabulary","terms":[{"code":"controlled-access-type-1","label":{"ja":"制限公開（Type I）","en":"Controlled-access (Type I)"}}]},{"key":"type-of-data","label":{"ja":"データの種類","en":"Type of data"},"type":"text","text":{"ja":"BBJ第1コホート1,026名のWGSデータ","en":"WGS for 1,026 individuals"}}],"experiments":[{"label":{"ja":"WGS","en":"WGS"},"values":[{"key":"materials-and-participants","label":{"ja":"材料と対象者","en":"Materials and participants"},"type":"text","text":{"ja":"2003年から2007年度にバイオバンク・ジャパンに登録された47疾患患者：1,026名","en":"1,026 individuals"}},{"key":"health-status","label":{"ja":"健康状態","en":"Health status"},"type":"vocabulary","terms":[{"code":"affected","label":{"ja":"罹患","en":"Affected"}}]},{"key":"subject-count","label":{"ja":"対象者数","en":"Subject count"},"type":"number","numbers":[{"value":1026,"unit":null,"high":null}]},{"key":"subject-count-type","label":{"ja":"対象者数の数え方","en":"Counted as"},"type":"vocabulary","terms":[{"code":"individual","label":{"ja":"人数","en":"Individual"}}]},{"key":"cohort","label":{"ja":"コホート","en":"Cohort"},"type":"vocabulary","terms":[{"code":"biobank-japan","label":{"en":"BioBank Japan"}}]},{"key":"sample-description","label":{"ja":"試料説明","en":"Sample description"},"type":"text","text":{"ja":"末梢血から抽出したgDNA","en":"DNA extracted from peripheral blood cells"}},{"key":"tissue","label":{"ja":"組織","en":"Tissue"},"type":"vocabulary","terms":[{"code":"peripheral-blood","label":{"ja":"末梢血","en":"Peripheral blood"}}]},{"key":"sample-provider","label":{"ja":"試料の入手元","en":"Sample provider"},"type":"text","text":{"ja":null,"en":null}},{"key":"experimental-method","label":{"ja":"実験方法","en":"Experimental method"},"type":"vocabulary","terms":[{"code":"wgs","label":{"en":"WGS"}}]},{"key":"targets","label":{"ja":"測定対象","en":"Target"},"type":"text","text":{"ja":null,"en":null}},{"key":"reagents","label":{"ja":"試薬キット","en":"Reagent kit"},"type":"vocabulary","terms":[{"code":"truseq-nano-dna-library-prep-kit","label":{"en":"TruSeq Nano DNA Library Prep Kit"}}]},{"key":"fragmentation","label":{"ja":"断片化","en":"Fragmentation"},"type":"text","text":{"ja":"超音波断片化","en":"Ultrasonic fragmentation"}},{"key":"platform","label":{"ja":"プラットフォーム","en":"Platform"},"type":"vocabulary","terms":[{"code":"illumina-hiseq-2500","label":{"en":"Illumina HiSeq 2500"}}]},{"key":"read-type","label":{"ja":"リードタイプ","en":"Read type"},"type":"vocabulary","terms":[{"code":"paired-end","label":{"ja":"ペアエンド","en":"Paired-end"}}]},{"key":"read-length","label":{"ja":"リード長","en":"Read length"},"type":"number","numbers":[{"value":160,"unit":"bp","high":null}]},{"key":"reference-sequence","label":{"ja":"参照ゲノム","en":"Reference genome"},"type":"vocabulary","terms":[{"code":"grch37","label":{"en":"GRCh37"}}]},{"key":"mapping","label":{"ja":"マッピング","en":"Mapping"},"type":"text","text":{"ja":"BWA mem 0.7.12","en":"BWA mem 0.7.12"}},{"key":"read-deduplication","label":{"ja":"リード重複除去","en":"Read deduplication"},"type":"text","text":{"ja":"Picard 2.10.6","en":"Picard 2.10.6"}},{"key":"calibration-for-re-alignment-and-base-quality","label":{"ja":"再アライメント・塩基品質補正","en":"Realignment and base quality recalibration"},"type":"text","text":{"ja":"GATK 3.7","en":"GATK 3.7"}},{"key":"mapping-quality","label":{"ja":"マッピング品質","en":"Mapping quality"},"type":"text","text":{"ja":"GATK 3.7 HaplotypeCallerで変異コール時にMAPQ<20のリードを除外","en":"Reads with MAPQ<20 were excluded at variant calling with GATK 3.7 HaplotypeCaller"}},{"key":"filtering","label":{"ja":"QC・フィルタリング","en":"QC and filtering"},"type":"text","text":{"ja":"リードのbase qualityが全体的に悪い検体、リード毎の%GCの結果にて異常を示した検体を除去。\nAlignment後、Low mapping rate検体、Insert sizeがおかしい検体、メタデータの性別情報とalingment結果より推定される性別情報が不一致な検体、性染色体異常疑いの検体を除去。\nGenotyping時に、VQSR、DP/GP filter (DP < 5, GQ < 20, DP > 60 && GQ < 95を除去)、heterozygosity filter (F>=0.05 を除去)、HWE filter (p < 10-6を除去)、Repeat & Low Complexity filterを実施。\n1000 genomes projectと合わせたPCAを実施し、日本人クラスタから大きく外れる検体を除外。\nその後、Genome-In-A-Bottleプロジェクトから公開されているHighConfidenceRegionリストに記載のある領域のバリアントにフラグを付与。","en":"Data with bad base quality and high %GC content were removed.\nAlignment:\nData matched for the following conditions were removed.\n- Low mapping rate\n- Different insert size\n- Gender information mismatch between meta-data and genotype data\n- Suspected sex chromosome aberration\nGenotyping:\nGATK's best practices include a variant filtering step following Variant Quality Score Recalibration (VQSR)\n- DP/GP (DP < 5, GQ < 20, DP > 60, GQ < 95)\n- Heterozygosity (F>=0.05)\n- Hardy-Weinberg equilibrium (p < 10^-6)\n- Repeat & Low Complexity\nPrincipal Component Analysis (PCA):\nPCA was performed with individuals included in the 1000 genomes project and outliers from Japanese cluster were removed.\n\nAfter these filtering steps, variants located in the regions listed as the HighConfidenceRegion (Genome-In-A-Bottle project) were flagged."}},{"key":"analysis-methods","label":{"ja":"解析方法","en":"Analysis method"},"type":"text","text":{"ja":"GATK 3.7 HaplotypeCaller","en":"GATK 3.7 HaplotypeCaller"}},{"key":"coverage-depth","label":{"ja":"カバレッジ (深度)","en":"Coverage (depth)"},"type":"number","numbers":[{"value":31.8,"unit":"x","high":null}]},{"key":"variant-number","label":{"ja":"バリアント数","en":"Variant count"},"type":"number","numbers":[{"value":10202908,"unit":null,"high":null,"prefix":{"ja":"常染色体: ","en":"Autosomes: "}},{"value":410435,"unit":null,"high":null,"prefix":{"ja":"X染色体: ","en":"X chromosome: "}},{"value":76768387,"unit":null,"high":null,"prefix":{"ja":"常染色体: ","en":"Autosomes: "}},{"value":2898518,"unit":null,"high":null,"prefix":{"ja":"X染色体: ","en":"X chromosome: "}}]},{"key":"processed-data-type","label":{"ja":"加工データの種類","en":"Processed data type"},"type":"vocabulary","terms":[{"code":"alignment","label":{"ja":"アライメント","en":"Alignment"}},{"code":"variant-calls","label":{"ja":"変異コール","en":"Variant calls"}}]},{"key":"processed-data-dataset-id","label":{"ja":"加工データのデータセットID","en":"Dataset ID of processed data"},"type":"text","text":{"ja":"JGAD000690\nJGAD000758（joint call）","en":"JGAD000690\nJGAD000758 (joint call)"}},{"key":"processing-method","label":{"ja":"加工方法","en":"Processing method"},"type":"text","text":{"ja":"加工済みデータ一覧（Whole Genome Sequencing解析）","en":"List of processed data (Whole Genome Sequencing Analysis)"}},{"key":"policies","label":{"ja":"利用ポリシー","en":"Data use policy"},"type":"vocabulary","terms":[{"code":"nbdc-data-sharing-policy-jgap000001","label":{"ja":"NBDC データ共有ポリシー (JGAP000001)","en":"NBDC data sharing policy (JGAP000001)"}}]}]},{"label":{"ja":"WGS リファレンスパネル (autosomes + X)","en":"WGS reference panel (autosomes + X)"},"values":[{"key":"materials-and-participants","label":{"ja":"材料と対象者","en":"Materials and participants"},"type":"text","text":{"ja":"・1,026名のWGSデータ（JGAD000220）を含むバイオバンクジャパン1,037名のWGSデータのvcfファイル（WGSデータの分子データは上記参照のこと）\n・1000ゲノムプロジェクト（Phase3v5）2,504名のWGSのvcfファイル（ftp://ftp.1000genomes.ebi.ac.uk/vol1/ftp/release/20130502/）","en":"- WGS data (JGAD000220) of the biobank Japan project (N=1,037)\n- WGS data of 1KGP p3v5 ALL (N=2,504) (ftp://ftp.1000genomes.ebi.ac.uk/vol1/ftp/release/20130502/)"}},{"key":"subject-count","label":{"ja":"対象者数","en":"Subject count"},"type":"number","numbers":[{"value":1037,"unit":null,"high":null}]},{"key":"subject-count-type","label":{"ja":"対象者数の数え方","en":"Counted as"},"type":"vocabulary","terms":[{"code":"individual","label":{"ja":"人数","en":"Individual"}}]},{"key":"cohort","label":{"ja":"コホート","en":"Cohort"},"type":"vocabulary","terms":[{"code":"biobank-japan","label":{"en":"BioBank Japan"}},{"code":"1000-genomes-project","label":{"en":"1000 Genomes Project"}}]},{"key":"experimental-method","label":{"ja":"実験方法","en":"Experimental method"},"type":"vocabulary","terms":[{"code":"wgs","label":{"en":"WGS"}}]},{"key":"targets","label":{"ja":"測定対象","en":"Target"},"type":"text","text":{"ja":null,"en":null}},{"key":"reference-sequence","label":{"ja":"参照ゲノム","en":"Reference genome"},"type":"vocabulary","terms":[{"code":"grch37","label":{"en":"GRCh37"}}]},{"key":"mapping","label":{"ja":"マッピング","en":"Mapping"},"type":"text","text":{"ja":"BWA-MEM (version 0.7.5a)","en":"BWA-MEM (version 0.7.5a)"}},{"key":"read-deduplication","label":{"ja":"リード重複除去","en":"Read deduplication"},"type":"text","text":{"ja":"picard (versions 1.106)","en":"picard (versions 1.106)"}},{"key":"calibration-for-re-alignment-and-base-quality","label":{"ja":"再アライメント・塩基品質補正","en":"Realignment and base quality recalibration"},"type":"text","text":{"ja":"GATK (version 3.2-2)","en":"GATK (ver.3.2-2)"}},{"key":"mapping-quality","label":{"ja":"マッピング品質","en":"Mapping quality"},"type":"text","text":{"ja":"MAPQ < 20 を除外（HaplotypeCaller）","en":"MAPQ < 20 were excluded (HaplotypeCaller)"}},{"key":"filtering","label":{"ja":"QC・フィルタリング","en":"QC and filtering"},"type":"text","text":{"ja":"We set exclusion criteria for genotypes as follows:\n(1) DP < 5, (2) GQ < 20, or (3) DP > 60 and GQ < 95, and regarded these genotypes as missing.\nVariants with call rates < 90% were excluded before variant quality score recalibration (VQSR).\nAfter VQSR, we excluded variants located in low-complexity regions (LCR), as defined by mdust software were excluded.\nFinally, we used BEAGLE to impute missing genotypes.","en":"We set exclusion criteria for genotypes as follows:\n(1) DP < 5, (2) GQ < 20, or (3) DP > 60 and GQ < 95, and regarded these genotypes as missing.\nVariants with call rates < 90% were excluded before variant quality score recalibration (VQSR).\nAfter VQSR, we excluded variants located in low-complexity regions (LCR), as defined by mdust software were excluded.\nFinally, we used BEAGLE to impute missing genotypes."}},{"key":"analysis-methods","label":{"ja":"解析方法","en":"Analysis method"},"type":"text","text":{"ja":"GATK HaplotypeCaller (version 3.2-2)","en":"GATK HaplotypeCaller (version 3.2-2)"}},{"key":"coverage-depth","label":{"ja":"カバレッジ (深度)","en":"Coverage (depth)"},"type":"number","numbers":[{"value":30,"unit":"x","high":null,"suffix":{"ja":" (目標深度)","en":" (aimed at depth)"}}]},{"key":"variant-number","label":{"ja":"バリアント数","en":"Variant count"},"type":"number","numbers":[{"value":61608817,"unit":null,"high":null,"suffix":{"ja":" variants","en":" variants"}},{"value":59387070,"unit":null,"high":null,"prefix":{"ja":"常染色体: ","en":"Autosomes: "},"suffix":{"ja":" variants","en":" variants"}},{"value":2221747,"unit":null,"high":null,"prefix":{"ja":"X染色体: ","en":"X chromosome: "},"suffix":{"ja":" variants","en":" variants"}}]},{"key":"processed-data-type","label":{"ja":"加工データの種類","en":"Processed data type"},"type":"vocabulary","terms":[{"code":"reference-panel","label":{"ja":"リファレンスパネル","en":"Reference panel"}},{"code":"variant-calls","label":{"ja":"変異コール","en":"Variant calls"}}]},{"key":"processed-data-dataset-id","label":{"ja":"加工データのデータセットID","en":"Dataset ID of processed data"},"type":"text","text":{"ja":"JGAD000679","en":"JGAD000679"}},{"key":"processing-method","label":{"ja":"加工方法","en":"Processing method"},"type":"text","text":{"ja":"加工済みデータ一覧（Imputation reference）","en":"List of processed data (Imputation reference)"}},{"key":"policies","label":{"ja":"利用ポリシー","en":"Data use policy"},"type":"vocabulary","terms":[{"code":"nbdc-data-sharing-policy-jgap000001","label":{"ja":"NBDC データ共有ポリシー (JGAP000001)","en":"NBDC data sharing policy (JGAP000001)"}}]}]},{"label":{"ja":"WGS インピュテーションパネル (autosomes + X)","en":"WGS imputation panel (autosomes + X)"},"values":[{"key":"materials-and-participants","label":{"ja":"材料と対象者","en":"Materials and participants"},"type":"text","text":{"ja":"【JGAD000867】\nバイオバンクジャパン1,026名のWGSデータ（JGAD000220）のvcfファイル（WGSデータの分子データは上記参照のこと）\n【JGAD000868】\nバイオバンクジャパン1,964名のWGSデータ（JGAD000495）のvcfファイル（WGSデータの分子データは上記参照のこと）","en":"[JGAD000867]\n- WGS data (JGAD000220) of the biobank Japan project (N=1,026)\n[JGAD000868]\n- WGS data (JGAD000495) of the biobank Japan project (N=1,964)"}},{"key":"subject-count","label":{"ja":"対象者数","en":"Subject count"},"type":"number","numbers":[{"value":2990,"unit":null,"high":null}]},{"key":"subject-count-type","label":{"ja":"対象者数の数え方","en":"Counted as"},"type":"vocabulary","terms":[{"code":"individual","label":{"ja":"人数","en":"Individual"}}]},{"key":"cohort","label":{"ja":"コホート","en":"Cohort"},"type":"vocabulary","terms":[{"code":"biobank-japan","label":{"en":"BioBank Japan"}}]},{"key":"experimental-method","label":{"ja":"実験方法","en":"Experimental method"},"type":"vocabulary","terms":[{"code":"wgs","label":{"en":"WGS"}}]},{"key":"targets","label":{"ja":"測定対象","en":"Target"},"type":"text","text":{"ja":null,"en":null}},{"key":"reference-sequence","label":{"ja":"参照ゲノム","en":"Reference genome"},"type":"vocabulary","terms":[{"code":"grch38","label":{"en":"GRCh38"}}]},{"key":"mapping","label":{"ja":"マッピング","en":"Mapping"},"type":"text","text":{"ja":"bwa mem（version 0.7.15）","en":"bwa mem (version 0.7.15)"}},{"key":"read-deduplication","label":{"ja":"リード重複除去","en":"Read deduplication"},"type":"text","text":{"ja":"GATK MarkDuplicates（version 4.1.0.0）","en":"GATK MarkDuplicates (version 4.1.0.0)"}},{"key":"calibration-for-re-alignment-and-base-quality","label":{"ja":"再アライメント・塩基品質補正","en":"Realignment and base quality recalibration"},"type":"text","text":{"ja":null,"en":null}},{"key":"mapping-quality","label":{"ja":"マッピング品質","en":"Mapping quality"},"type":"text","text":{"ja":null,"en":null}},{"key":"filtering","label":{"ja":"QC・フィルタリング","en":"QC and filtering"},"type":"text","text":{"ja":"生殖系列（germline）の全ゲノムシークエンスデータの加工を行い、aggregate VCFを計算した。その後、下記の条件でvariantsのフィルタリングを行った。\n(1) VQSRフィルタを通過しなかったvariantsの除外\n(2) Multi-allelic sitesの除外\n(3) Call rate が低い（95%未満）variantsの除外\n(4) Hardy-Weinberg平衡から逸脱している（P<1e-10）variantsの除外\n(5) Minor allele count（MAC）が小さい（< 2）variantsの除外","en":"Germline whole genome sequencing data were processed, and the aggregate VCF was calculated. Variants were then filtered based on the following conditions:\n(1) Variants that did not pass the VQSR filter were excluded\n(2) Multi-allelic sites were excluded\n(3) Variants with a call rate below 95% were excluded\n(4) Variants deviating from Hardy-Weinberg equilibrium (P < 1e-10) were excluded\n(5) Variants with a minor allele count (MAC) less than 2 were excluded"}},{"key":"analysis-methods","label":{"ja":"解析方法","en":"Analysis method"},"type":"text","text":{"ja":"GATK HaplotypeCaller -ERC GVCF（version 4.1.0.0）\nバリアント検出時のパラメータである ploidy は次のように設定した。\n常染色体および pseudoautosomal region（PAR） 領域：ploidy=2\nX染色体のnon-PAR領域：ploidy=2（女性） および ploidy=1（男性）\nY染色体の non-PAR 領域：ploidy=1（男性）","en":"GATK HaplotypeCaller -ERC GVCF (version 4.1.0.0)\nThe ploidy for variant call was set as follows:\nAutosomes and pseudoautosomal regions (PARs): ploidy=2\nNon-PARs on the X chromosome: ploidy=2 (female) and ploidy=1 (male)\nNon-PARs on the Y chromosome: ploidy=1 (male)"}},{"key":"processed-data-type","label":{"ja":"加工データの種類","en":"Processed data type"},"type":"vocabulary","terms":[{"code":"reference-panel","label":{"ja":"リファレンスパネル","en":"Reference panel"}},{"code":"variant-calls","label":{"ja":"変異コール","en":"Variant calls"}}]},{"key":"processed-data-dataset-id","label":{"ja":"加工データのデータセットID","en":"Dataset ID of processed data"},"type":"text","text":{"ja":"JGAD000867\nJGAD000868","en":"JGAD000867\nJGAD000868"}},{"key":"processing-method","label":{"ja":"加工方法","en":"Processing method"},"type":"text","text":{"ja":"加工済みデータ一覧（Imputation reference）","en":"List of processed data (Imputation reference)"}},{"key":"policies","label":{"ja":"利用ポリシー","en":"Data use policy"},"type":"vocabulary","terms":[{"code":"nbdc-data-sharing-policy-jgap000001","label":{"ja":"NBDC データ共有ポリシー (JGAP000001)","en":"NBDC data sharing policy (JGAP000001)"}}]}]}],"dataVolume":75142080702713,"fileFormats":[{"code":"fastq","label":{"en":"FASTQ"}},{"code":"vcf","label":{"en":"VCF"}},{"code":"html","label":{"en":"HTML"}}]}