{"id":"JGAD000873","research":"hum0014","url":"https://humandbs.dbcls.jp/dataset/JGAD000873","datePublished":"2024-11-29","dateModified":"2024-11-29","values":[{"key":"access-criteria","label":{"ja":"アクセス制限","en":"Access criteria"},"type":"vocabulary","terms":[{"code":"controlled-access-type-1","label":{"ja":"制限公開（Type I）","en":"Controlled-access (Type I)"}}]},{"key":"type-of-data","label":{"ja":"データの種類","en":"Type of data"},"type":"text","text":{"ja":"バイオバンクジャパン7,472名のWGSデータと1000ゲノムプロジェクト（Phase3v5）2,504名のWGSのvcfファイルを統合したgenotype imputation用のreference panel","en":"Imputation reference panel for 7,472 Japanese WGS and 2,504 1000 Genome Project data"}}],"experiments":[{"label":{"ja":"WGS + リファレンスパネル (autosomes)","en":"WGS + reference panel (autosomes)"},"values":[{"key":"materials-and-participants","label":{"ja":"材料と対象者","en":"Materials and participants"},"type":"text","text":{"ja":"【WGS】\nBBJ第1コホート：7,472名\n15-30×深度：3,256名、3×深度：4,216名\n（下記データを含む）\n・2003年から2007年度にバイオバンク・ジャパンに登録された47疾患患者 1,026名（JGAD000220）\n・バイオバンク・ジャパン第1コホートに含まれる胃がん 256名（JGAD000831）\n・2003年から2017年度にバイオバンク・ジャパンに登録された症例 1,964名（JGAD000495）\n・バイオバンク・ジャパン第1コホートに含まれる糖尿病のうち 2,157名（JGAD000833）\n・バイオバンク・ジャパン第1コホートに含まれる胃がんのうち 2,059名（JGAD000834）\n【reference panel】\n・BBJ第1コホート7,472名のWGSのvcfファイル\n・1000ゲノムプロジェクト（Phase3v5）2,504名のWGSのvcfファイル（ftp://ftp.1000genomes.ebi.ac.uk/vol1/ftp/release/20130502/）","en":"[WGS]\nSubjects registered in BioBank Japan: 7,472 individuals\n15-30x depth: 3,256 individuals, 3x depth: 4,216 individuals\nincluding followed data\n- 1,026 individuals (JGAD000220)\n- 256 gastric cancer patients registered in BBJ 1st cohort (JGAD000831)\n- 1,765 myocardial infarction patients and 199 dementia patients (JGAD000495)\n- 2,157 diabetes patients registered in BBJ 1st cohort (JGAD000833)\n- 2,059 gastric cancer patients registered in BBJ 1st cohort (JGAD000834)\n[reference panel]\n- WGS data of the BBJ 1st cohort (N=7,472)\n- WGS data of 1KGP p3v5 ALL (N=2,504) (ftp://ftp.1000genomes.ebi.ac.uk/vol1/ftp/release/20130502/)"}},{"key":"health-status","label":{"ja":"健康状態","en":"Health status"},"type":"vocabulary","terms":[{"code":"affected","label":{"ja":"罹患","en":"Affected"}}]},{"key":"subject-count","label":{"ja":"対象者数","en":"Subject count"},"type":"number","numbers":[{"value":7472,"unit":null,"high":null}]},{"key":"subject-count-type","label":{"ja":"対象者数の数え方","en":"Counted as"},"type":"vocabulary","terms":[{"code":"individual","label":{"ja":"人数","en":"Individual"}}]},{"key":"cohort","label":{"ja":"コホート","en":"Cohort"},"type":"vocabulary","terms":[{"code":"biobank-japan","label":{"en":"BioBank Japan"}},{"code":"1000-genomes-project","label":{"en":"1000 Genomes Project"}}]},{"key":"population","label":{"ja":"対象集団","en":"Population"},"type":"vocabulary","terms":[{"code":"japanese","label":{"ja":"日本人","en":"Japanese"}}]},{"key":"sample-description","label":{"ja":"試料説明","en":"Sample description"},"type":"text","text":{"ja":"末梢血から抽出したgDNA","en":"DNA extracted from peripheral blood cells"}},{"key":"tissue","label":{"ja":"組織","en":"Tissue"},"type":"vocabulary","terms":[{"code":"peripheral-blood","label":{"ja":"末梢血","en":"Peripheral blood"}}]},{"key":"is-tumor","label":{"ja":"腫瘍 / 非腫瘍","en":"Tumor / normal"},"type":"vocabulary","terms":[{"code":"normal","label":{"ja":"非腫瘍","en":"Normal"}}]},{"key":"sample-provider","label":{"ja":"試料の入手元","en":"Sample provider"},"type":"text","text":{"ja":null,"en":null}},{"key":"experimental-method","label":{"ja":"実験方法","en":"Experimental method"},"type":"vocabulary","terms":[{"code":"wgs","label":{"en":"WGS"}}]},{"key":"targets","label":{"ja":"測定対象","en":"Target"},"type":"text","text":{"ja":null,"en":null}},{"key":"reagents","label":{"ja":"試薬キット","en":"Reagent kit"},"type":"vocabulary","terms":[{"code":"truseq-nano-dna-library-prep-kit","label":{"en":"TruSeq Nano DNA Library Prep Kit"}}]},{"key":"fragmentation","label":{"ja":"断片化","en":"Fragmentation"},"type":"text","text":{"ja":"超音波断片化","en":"Ultrasonic fragmentation"}},{"key":"platform","label":{"ja":"プラットフォーム","en":"Platform"},"type":"vocabulary","terms":[{"code":"illumina-hiseq-2500","label":{"en":"Illumina HiSeq 2500"}},{"code":"illumina-hiseq-x","label":{"en":"Illumina HiSeq X"}}]},{"key":"read-type","label":{"ja":"リードタイプ","en":"Read type"},"type":"vocabulary","terms":[{"code":"paired-end","label":{"ja":"ペアエンド","en":"Paired-end"}}]},{"key":"read-length","label":{"ja":"リード長","en":"Read length"},"type":"number","numbers":[{"value":125,"unit":"bp","high":null},{"value":126,"unit":"bp","high":null},{"value":151,"unit":"bp","high":null},{"value":161,"unit":"bp","high":null}]},{"key":"reference-sequence","label":{"ja":"参照ゲノム","en":"Reference genome"},"type":"vocabulary","terms":[{"code":"grch37","label":{"en":"GRCh37"}}]},{"key":"mapping","label":{"ja":"マッピング","en":"Mapping"},"type":"text","text":{"ja":"15-30×深度：BWA mem 0.7.12\n3×深度：BWA mem 0.7.5a","en":"15-30x depth: BWA mem 0.7.12\n3x depth: BWA mem 0.7.5a"}},{"key":"read-deduplication","label":{"ja":"リード重複除去","en":"Read deduplication"},"type":"text","text":{"ja":"15-30×深度：Picard 2.10.6\n3×深度：Picard 1.106、2.5.0","en":"15-30x depth: Picard 2.10.6\n3x depth: Picard 1.106, 2.5.0"}},{"key":"calibration-for-re-alignment-and-base-quality","label":{"ja":"再アライメント・塩基品質補正","en":"Realignment and base quality recalibration"},"type":"text","text":{"ja":"GATK v.3.2-2","en":"GATK v.3.2-2"}},{"key":"mapping-quality","label":{"ja":"マッピング品質","en":"Mapping quality"},"type":"text","text":{"ja":"15-30×深度：GATK 3.7 HaplotypeCallerで変異コール時にMAPQ<20 のリードを除外\n3×深度：GotCloudで変異コール時にMAPQ<20のリードを除外","en":"15-30x depth: Reads with MAPQ<20 were excluded at variant calling with GATK 3.7 HaplotypeCaller\n3x depth: Reads with MAPQ<20 were excluded at variant calling with GotCloud"}},{"key":"filtering","label":{"ja":"QC・フィルタリング","en":"QC and filtering"},"type":"text","text":{"ja":"We set exclusion criteria for genotypes sequenced at high depth (30x and 15x) as follows:\n(1) DP < 5, (2) GQ < 20, or (3) DP > 60 and GQ < 95, and regarded these genotypes as missing.\nVariants with call rates < 90% were excluded before variant quality score recalibration (VQSR).\nAfter VQSR, variants located in low-complexity regions (LCR), as defined by mdust software, were excluded in the high depth (30x) dataset.\nFinally, we used BEAGLE to impute missing genotypes.","en":"We set exclusion criteria for genotypes sequenced at high depth (30x and 15x) as follows:\n(1) DP < 5, (2) GQ < 20, or (3) DP > 60 and GQ < 95, and regarded these genotypes as missing.\nVariants with call rates < 90% were excluded before variant quality score recalibration (VQSR).\nAfter VQSR, variants located in low-complexity regions (LCR), as defined by mdust software, were excluded in the high depth (30x) dataset.\nFinally, we used BEAGLE to impute missing genotypes."}},{"key":"analysis-methods","label":{"ja":"解析方法","en":"Analysis method"},"type":"text","text":{"ja":"15-30×深度：GATK 3.7 HaplotypeCaller\n3×深度：GotCloud v1.17.5","en":"15-30x depth: GATK 3.7 HaplotypeCaller\n3x depth: GotCloud v1.17.5"}},{"key":"coverage-depth","label":{"ja":"カバレッジ (深度)","en":"Coverage (depth)"},"type":"number","numbers":[{"value":30,"unit":"x","high":null},{"value":15,"unit":"x","high":null},{"value":3,"unit":"x","high":null}]},{"key":"variant-number","label":{"ja":"バリアント数","en":"Variant count"},"type":"number","numbers":[{"value":4535276,"unit":null,"high":null},{"value":80753886,"unit":null,"high":null},{"value":85328475,"unit":null,"high":null,"suffix":{"ja":" variants","en":" variants"}}]},{"key":"policies","label":{"ja":"利用ポリシー","en":"Data use policy"},"type":"vocabulary","terms":[{"code":"nbdc-data-sharing-policy-jgap000001","label":{"ja":"NBDC データ共有ポリシー (JGAP000001)","en":"NBDC data sharing policy (JGAP000001)"}}]}]}],"dataVolume":40455354428,"fileFormats":[{"code":"vcf","label":{"en":"VCF"}}]}