Skip to content
NBDC Human Database

No datasets in the cart.

Due to system maintenance, the application system, application review by the Data Access Committee will be unavailable during the following period.
Schedule: October 5th (Mon), 2026, 9:00 - October 7th (Wed), 2026, 15:00 (JST)
We apologize for any inconvenience this may cause and appreciate your understanding.

We are currently receiving a large number of applications for data submission, and the review process is taking longer than usual.We sincerely apologize for the delay and kindly ask for your understanding. When submitting an application, we would greatly appreciate it if you could allow sufficient time for the processing.

Following a change to our organizational structure effective April 1, 2026, this division has been renamed from the "Database Center for Life Science, Joint Support-Center for Data Science Research" to the "Database Division for Life Science (DBCLS), BioData Science Initiative (BSI), National Institute of Genetics (NIG)". Where the former name still appears in the guidelines, please read it as the new name.

Dataset ID

JGAD000117

Type of data
NGS (WGS)
NGS (Exome)
Access criteria
Controlled-access (Type I)
Total data volume
3.0 TB
File formats
  • FASTQ
Research
hum0103
Date published
2020-09-28
Date modified
2021-11-26

Analysis method

WGS

Materials and participants
biliary tract cancer (ICD10: C22, 23, 24): 14 cases + 3 cases
cancer tissues: 23 samples + 6 samples
paired non-cancer tissues: 14 samples + 3 samples
  • Health status
    Mixed
  • Subject count
    17 (Individual)
Disease
biliary tract cancer (C22, C23, C24)
Sample description
DNAs extracted from cancer and paired non-cancer (normal) tissues from biliary tract cancer patients
  • Tumor / normal
    Mixed
Sample provider
N/A
Experimental method
WGS
Target
N/A
Reagent kit
TruSeq DNA Sample Prep Kit
Fragmentation
Ultrasonic fragmentation (Covaris)
Platform
Illumina HiSeq 2000
Illumina HiSeq 2500
Illumina NovaSeq 6000
Read type
Paired-end
Read length
100 bp
Reference genome
GRCh37
Mapping
BWA mem 0.7.12
Read deduplication
Picard 2.10.6
Realignment and base quality recalibration
GATK 3.7
Mapping quality
Reads with MAPQ< 20 were excluded at variant calling with GATK 3.7 HaplotypeCaller
QC and filtering
Data with bad base quality and high %GC content were removed.
Aligment:
Data matched for the following condition were removed.
- Low mapping rate
- Different insert size
- Gender information mismatch between meta-data and genotype data
- Suspected sex chromosome aberration
Genotyping:
GATK's best practices includes a variant filtering step following Variant Quality Score Recalibration (VQSR)
- DP/GP (DP < 5, GQ < 20, DP > 60, GQ < 95)
- Heterozygosity (F>=0.05)
- Hardy-Weinberg equilibrium (p < 10^-6)
- Repeat & Low Complexity
Principal Component Analysis (PCA):
PCA was performed with individuals included in the 1000 genomes project and outliers from Japanese cluster were removed.

After these filtering steps, variants located in the regions listed as the HighConfidenceRegion (Genome-In-A-Bottle project) were flagged.
Analysis method
GATK 3.7 HaplotypeCaller
Coverage (depth)
HiSeq 2000/2500: 31.8x
NovaSeq 6000: 28x
Variant count
Autosomes: 10,202,908
X chromosome: 410,435
Autosomes: 76,768,387
X chromosome: 2,898,518
European Genome-phenome Archive Accession
Included in EGAS00001000678 [EGAD00001000809]