Skip to content
NBDC Human Database

No datasets in the cart.

Due to system maintenance, the application system, application review by the Data Access Committee will be unavailable during the following period.
Schedule: October 5th (Mon), 2026, 9:00 - October 7th (Wed), 2026, 15:00 (JST)
We apologize for any inconvenience this may cause and appreciate your understanding.

We are currently receiving a large number of applications for data submission, and the review process is taking longer than usual.We sincerely apologize for the delay and kindly ask for your understanding. When submitting an application, we would greatly appreciate it if you could allow sufficient time for the processing.

Following a change to our organizational structure effective April 1, 2026, this division has been renamed from the "Database Center for Life Science, Joint Support-Center for Data Science Research" to the "Database Division for Life Science (DBCLS), BioData Science Initiative (BSI), National Institute of Genetics (NIG)". Where the former name still appears in the guidelines, please read it as the new name.

Dataset ID

JGAD001048

Type of data
NGS (WGS, RNA-seq)
Access criteria
Controlled-access (Type I)
Total data volume
4.6 TB
File formats
  • FASTQ
  • MAF
  • CSV
Research
hum0567
Date published
2026-09-24
Date modified
2026-09-24

Analysis method

WGS

Materials and participants
liver cancer caused by rare liver diseases, especially FALD (ICD10: C220): 13 cases
tumor tissue: 13 samples
non-tumor tissue: 13 samples
  • Health status
    Affected
  • Subject count
    13 (Individual)
  • Population
    Japanese
Disease
liver cancer caused by rare liver diseases, especially FALD (C220)
Sample description
DNAs extracted from tumor and non-tumor tissues
  • Tissue
    Liver
  • Tumor / normal
    Mixed
Sample provider
N/A
Experimental method
WGS
Target
N/A
Reagent kit
TruSeq DNA PCR-Free Library Prep Kit
Fragmentation
Ultrasonic fragmentation (Covaris)
Platform
Illumina NovaSeq 6000
Illumina NovaSeq X Plus
Read type
Paired-end
Read length
150 bp
Reference genome
GRCh38
Mapping
BWA-MEM
Read deduplication
Picard MarkDuplicates (Duplication rate: approx. 15%)
Realignment and base quality recalibration
GATK BQSR
Mapping quality
MAPQ >= 20
QC and filtering
- Adapter trimming via MarkIlluminaAdapters and SamToFastq with clipping parameters.
- Overlapping read clipping performed by MergeBamAlignment (-CLIP_OVERLAPPING_READS true).
- Duplicate read filtering using GATK MarkDuplicates.
- Base quality score recalibration (BQSR) against dbSNP and known indel resources.
- Artifact and contamination filtering using GATK FilterMutectCalls default pipeline.
- Filtered variant removal during functional annotation via Funcotator.
cf. Supplementary Methods of DOI: 10.1097/HEP.0000000000001693 for detailed pipeline configurations and parameters
Analysis method
GATK Mutect2 (v4.6.2.0) + Funcotator
Nextflow oncoanalyser (v2.1.0) / PURPLE
Nextflow oncoanalyser (v2.1.0) + R (StructuralVariantAnnotation / VariantAnnotation)
Coverage (depth)
100x
Variant count
1010 Indels (Median)
189–2786 Indels (Range)
64 SVs (Median)
2–160 SVs (Range)
5373 SNVs (Median)
2000–12,634 SNVs (Range)

RNA-seq

Materials and participants
liver cancer caused by rare liver diseases, especially FALD (ICD10: C220): 13 cases
tumor tissue: 13 samples
non-tumor tissue: 13 samples
  • Health status
    Affected
  • Subject count
    13 (Individual)
  • Population
    Japanese
Disease
liver cancer caused by rare liver diseases, especially FALD (C220)
Sample description
RNAs extracted from tumor and non-tumor tissues
  • Tissue
    Liver
  • Tumor / normal
    Mixed
Sample provider
N/A
Experimental method
RNA-seq
Target
N/A
Reagent kit
NEBNext Ultra II Directional RNA Library Prep Kit for Illumina
Fragmentation
Chemical fragmentation [divalent cations, heat treatment]
Platform
Illumina NovaSeq 6000
Illumina NovaSeq X Plus
Read type
Paired-end
Read length
150 bp
Reference genome
GRCh38
Mapping
STAR (v2.7.10a)
Mapping quality
>91%
QC and filtering
None (RSEM internally handles multi-mapping reads using EM algorithm)
Analysis method
STAR (v2.7.10a) + RSEM (v1.3.3)
Gene count
61,552