Skip to content
NBDC Human Database

No datasets in the cart.

Due to system maintenance, the application system, application review by the Data Access Committee will be unavailable during the following period.
Schedule: October 5th (Mon), 2026, 9:00 - October 7th (Wed), 2026, 15:00 (JST)
We apologize for any inconvenience this may cause and appreciate your understanding.

We are currently receiving a large number of applications for data submission, and the review process is taking longer than usual.We sincerely apologize for the delay and kindly ask for your understanding. When submitting an application, we would greatly appreciate it if you could allow sufficient time for the processing.

Following a change to our organizational structure effective April 1, 2026, this division has been renamed from the "Database Center for Life Science, Joint Support-Center for Data Science Research" to the "Database Division for Life Science (DBCLS), BioData Science Initiative (BSI), National Institute of Genetics (NIG)". Where the former name still appears in the guidelines, please read it as the new name.

Dataset ID

JGAD000404

Type of data
bam/gvcf data of NGS (WGS)
Access criteria
Controlled-access (Type I)
Total data volume
25.2 TB
File formats
  • BAM
  • BAI
  • VCF
  • TBI
Research
hum0158
Date published
2021-05-25
Date modified
2021-05-25

Analysis method

WGS

Materials and participants
liver cancer (ICD10: C220, 221, 227): 258 cases + 5 cases
cancer tissues: 301 samples + 5 samples
paired non-cancer tissues: 265 samples (257 blood samples, 3 liver tissues + 5 blood samples)
  • Health status
    Affected
  • Subject count
    263 (Individual)
Disease
hepatocellular carcinoma (C220)
intrahepatic cholangiocarcinoma (C221)
combined hepatocellular-cholangiocarcinoma (C227)
Sample description
DNAs extracted from cancer tissues and paired non-cancer tissues or blood samples from liver cancer patients
  • Tissue
    Peripheral blood
  • Tumor / normal
    Mixed
Sample provider
N/A
Experimental method
WGS
Target
N/A
Reagent kit
Paired-End DNA Sample Prep Kit
TruSeq DNA Sample Prep Kit
TruSeq Nano DNA Library Prep Kit
Fragmentation
Ultrasonic fragmentation (Covaris)
Platform
Illumina Genome Analyzer IIx
Illumina HiSeq 2000
Illumina NovaSeq 6000
Read type
Paired-end
Read length
100 bp
Reference genome
GRCh37
Mapping
BWA mem 0.7.12
Read deduplication
Picard 2.10.6
Realignment and base quality recalibration
GATK 3.7
Mapping quality
Reads with MAPQ<20 were excluded at variant calling with GATK 3.7 HaplotypeCaller
QC and filtering
Data with bad base quality and high %GC content were removed.
Aligment:
Data matched for the following condition were removed.
- Low mapping rate
- Different insert size
- Gender information mismatch between meta-data and genotype data
- Suspected sex chromosome aberration
Genotyping:
GATK's best practices includes a variant filtering step following Variant Quality Score Recalibration (VQSR)
- DP/GP (DP < 5, GQ < 20, DP > 60, GQ < 95)
- Heterozygosity (F>=0.05)
- Hardy-Weinberg equilibrium (p < 10^-6)
- Repeat & Low Complexity
Principal Component Analysis (PCA):
PCA was performed with individuals included in the 1000 genomes project and outliers from Japanese cluster were removed.

After these filtering steps, variants located in the regions listed as the HighConfidenceRegion (Genome-In-A-Bottle project) were flagged.
Analysis method
GATK 3.7 HaplotypeCaller
Coverage (depth)
HiSeq 2000: 31.8x
Genome Analyzer IIx: 30x
NovaSeq 6000: 28x
Variant count
Autosomes: 10,202,908
X chromosome: 410,435
Autosomes: 76,768,387
X chromosome: 2,898,518
Data summary
bam/vcf files of non-tumor tissues derived from 220 liver cancer patients