Skip to content
NBDC Human Database

No datasets in the cart.

Due to system maintenance, the application system, application review by the Data Access Committee will be unavailable during the following period.
Schedule: October 5th (Mon), 2026, 9:00 - October 7th (Wed), 2026, 15:00 (JST)
We apologize for any inconvenience this may cause and appreciate your understanding.

We are currently receiving a large number of applications for data submission, and the review process is taking longer than usual.We sincerely apologize for the delay and kindly ask for your understanding. When submitting an application, we would greatly appreciate it if you could allow sufficient time for the processing.

Following a change to our organizational structure effective April 1, 2026, this division has been renamed from the "Database Center for Life Science, Joint Support-Center for Data Science Research" to the "Database Division for Life Science (DBCLS), BioData Science Initiative (BSI), National Institute of Genetics (NIG)". Where the former name still appears in the guidelines, please read it as the new name.

Research ID

hum0331-v1Release info

Latest

Research title

Construction of control data for the promotion of genomic medicine for cancers and rare diseases

Research overview

Aims
Whole genome sequencing is being promoted for better medical care of rare diseases and cancers. For these disease genome analyses, whole genome sequencing (WGS) analysis data of the healthy control group is necessary. We conducted WGS analysis of healthy individuals for cancers and rare diseases from biobank specimens held by six National Centers (NCs) Biobanks in Japan, taking regional variations into consideration, and construct a genome database of healthy individuals and disease control groups.
Methods
DNA samples that meet the criteria for the study will be shipped from the biobank and subjected to WGS analysis at a contract analysis laboratory. WGS analysis will be performed on a Novaseq6000 sequencer using a PCR-free protocol to obtain a minimum output of 90 Gb. The read data in fastq format obtained from the analysis will be subjected to data analysis (mapping and variant calling) at the principal institute (National Center for Global Health and Medicine), and the data including variant information will be made into a database.
Participants/materials
A total of 9850 DNA samples from healthy individuals (including people with common complex diseases who do not have rare diseases or cancers) that can be utilized as controls for cancers and rare diseases studies. 560 were excluded after QC.

Datasets

CartDataset IDType of dataAnalysis methodAccess criteriaDate published
NHA000182NGS (WGS)
  • WGS
Unrestricted-access2023-02-01
NHA000181NGS (WGS)
  • WGS
Unrestricted-access2023-02-01

Data provider

Principal investigator
Katsushi Tokunaga
Affiliation
National Center for Global Health and Medicine Genome Medical Science

Research projects

NameURL
National Center Biobank Network
N/A

Grants

NameTitleProject number
Program for an Integrated Database of Clinical and Genomic Information, Japan Agency for Medical Research and Development (AMED)
Development of an integrated database of clinical genome information that contributes to the implementation of genomic medicine and the establishment of a continuous genomic medicine system in Japan
  • JP19kk0205012

Related publications

TitleDOIDataset ID
Exploring the genetic diversity of the Japanese population: Insights from a large-scale whole genome sequencing analysis