Main menu

Telomere-to-telomere sequencing (T2T) know-how document


Introduction

This know-how document provides supplementary information for the end-to-end workflow for telomere-to-telomere sequencing of the human genome using the Oxford Nanopore PromethION platform.

There are numerous routes for generating T2T and near-T2T de novo assemblies from long read data. The appropriate route depends on sample material, achievable read length and depth, compute resources, available phasing data and the intended analysis product. We recommend two protocols with associated analysis workflows for most users:



The scalable near-T2T workflow is designed for rapid generation of high-quality assemblies with some full-chromosome contigs from a single SQK-LSK114 ligation sequencing experiment. It targets approximately 40-45x coverage with a read N50 of around 30 kb, typically using around 1.5 PromethION flow cells. Without phasing data, it produces collapsed and dual (partially phased) assemblies; optional Hi-C or parental sequencing data can be added for haplotype-resolved assemblies.

The expert T2T workflow is designed to generate highly complete, haplotype-resolved diploid chromosome assemblies with accurate centromeres and other high-copy satellite regions. It combines SQK-ULK114 ultra-long sequencing (40–60x, read N50 ≥60 kb; optimally 80–100 kb) with SQK-LSK114 Pore-C for long-range phasing. It uses two PromethION flow cells for ultra-long data and one for Pore-C.

Scalable Assembly Expert Assembly
Main goal Some full chromosome contigs Many/near-complete diploid chromosomes with accurate centromeres
Best for Rapid, lower-input, lower-compute assembly Highest contiguity and structural accuracy
Sample 1 ml blood (fresh or frozen) 8.2 ml of blood (fresh)
Library preparation LSK (N50: 30 kb optimal*) ULK (N50: 60+ kb) +
Pore-C (requires LSK)
Phasing Optional: HiC or paternal data Built-in using Pore-C
Flow cells Approximately 1.5 flow cells for 40x coverage
Basecalling SUP**
3 flow cells (40X ULK + 10x Pore-C)
Basecalling: SUP**
Command line analysis Hifiasm (ONT) Dorado correct + Verkko
Expected output Collapsed/dual assembly unless phasing data added Haplotype-resolved assembly

*N50 of 20-50kb is suitable, less contiguous assemblies can be obtained with shorter reads, proportionally.
**HAC is possible but achieves slightly lower assembly quality and QV
***QV = Assembly quality at base level


The consumables and input requirements for each of the protocols are:

Preparation Preparation Kit R10.4.1 PromethION Flow Cells Input (whole blood)
Scalable Long read DNA sequencing SQK-LSK114 1–1.5 ≥ 1 ml
Expert Ultra-long DNA sequencing

Pore-C
SQK-ULK114

SQK-LSK114
2

1
2 x 1.6 ml (2 x ~6 million cells)

5–10 ml (~10 million cells)

Data and results

Scalable near-T2T

This rapid, lower-input method using standard long-read sequencing combined with assembly using Hifiasm is described in the Nature paper by Cheng et al (2026) Efficient near-telomere-to-telomere assembly of nanopore simplex reads. Briefly, the authors sequenced seven Genome in a Bottle (GIAB) samples, producing 9-20 gapless T2Ts, with contig N50s ranging from 88.3 – 135.7 Mb. Where parental data were available, trio-binning improved the assemblies by 2 to 3 additional gapless T2Ts.

Expert T2T

Six human samples were sequenced using the expert T2T method, combining ultra-long DNA sequencing with Pore-C data. All samples gave highly complete and contiguous assemblies, with 21 to 36 chromosomes assembled as gapless T2Ts and contigs ranging from 135.5 Mb to 154.5 Mb (Table 1). Figure 1 shows the ideogram of one sample, HG002, illustrating that most of the remaining gaps are in centromeres or acrocentric DNA. Further details and examples of how the data generated using expert T2T sequencing can be used to understand complex variant analysis can be found on our London Calling 2026 poster.


Sample Assembly size (Gb) Scaffold N50 (Mb) Contig N50 (Mb) Gapless chromosomes
HG002 6.04 146.8 135.9 30
GM24510 6.09 154.4 154.4 36
GM03786 6.08 154.5 137.3 21
GM21883 6.00 145.2 135.3 24
GM20291 6.04 146.2 143.1 32
GM20459 6.10 154.0 134.4 26

Table 1. Contiguity and number of gapless T2T chromosomes achieved for six samples sequenced using the expert T2T method.


T2T_Know_how_figure_1_Ideogram_of_hg002

Figure 1. Ideogram of HG002 assembly aligned to the Q100 T2T. This sample achieved 30 gapless chromosomes and a contig N50 of 135.9Mb.

Sequencing depth vs assembly performance

Assembly quality is expected to depend on a variety of factors, including sequencing depth and read length distribution as well as sample-specific characteristics. All else being equal, higher depths are expected to lead to better assemblies (Figure 2), up to 45X for scalable assembly (after which contiguity can actually get worse with the current (v0.25.0) version of Hifiasm) and 60X for the expert assembly (after which there are often no more non-rRNA gaps). For a given depth, longer reads are expected to yield better assemblies, as is a tighter read-length distribution. This is illustrated in the graph below, which shows performance vs. depth for a variety of datasets with each method. Data for the scalable assemblies for HG001, HG002, and HG005 are available from our GIAB dataset release.


T2T_Know_how_figure_2_assembly_statistics

Figure 2. Assembly statistics (# gapless chromosomes and contig N50) vs. sequencing depth for a diversity of samples/read lengths.

In addition to a generally higher contiguity, the expert method is expected to yield higher structural accuracy in high-copy satellite regions (predominantly centromeric and acrocentric satellites). Using the Q100 v1.1 HG002 reference for comparison, we see 58 large (>=50bp) indel errors in the 45X platinum method assembly (including 28 in centromeric/acrocentric satellites), vs 213 errors in the 45X scalable method assembly (including 149 in centromeric/acrocentric satellites).

Analysis

Current data processing requirements dictate that experienced bioinformaticians are needed, with an appropriate level of compute power as described in the community protocols. Details of how to undertake the analysis are also contained within the community protocol with appropriate Github links.

Use of the methods with non-human samples

The analysis methods used in both the scalable near-T2T and expert T2T methods have been described in the protocols for use on human samples. Both Hifiasm and Verkko expect to be able to reconstruct one or two haplotypes per chromosome, and are highly tailored to this use case.

Non-human samples which come from a single, eukaryotic, haploid or diploid individual should yield good results from these analytical methods. Note however, that laboratory methods may need adjusting to ensure sufficient coverage for the genome size. For bacterial isolates or metagenomic sequencing data, we recommend a dedicated tool like (meta)Flye, Myloasm, or NanoMDBG.

Samples which are not haploid or diploid, e.g. a polyploid plant or a non-karyotypical cell line (HeLa, Cho, HEK, tumor-derived lines, etc.) do not meet this criterion. Therefore, haplotype reconstructions are often erroneous to non-sensical, depending on the degree of ploidy/ploidy-variability. We have seen that Hifiasm is often capable of producing a reasonable collapsed assembly in these cases, though additional processing to remove redundant but lower-homology contigs may still be needed (e.g. the “purge_haplotigs” or “purge_dups” tools).

Information on genome assemblies

Basic terminology

A reconstruction of an organism’s genomic sequence is called a genome assembly. A contiguous stretch of assembled sequence, where we know every single letter and their relative orders, is known as a contig. Sometimes we know two contigs are adjacent in a genome, and we know their relative orientation, but we are missing information on the exact sequence in between – i.e. we have a gap. In this case we may choose to represent this in our genome assembly outputs as one contig followed by a series of ambiguous bases (“N” by convention) and then the next contig – this series of ordered and oriented contigs with gaps is known as a scaffold. If a contig covers a whole chromosome, spanning from one telomere to the other, and has sequence corresponding to a specific physical chromosome in a sample (i.e. is phased (see below) or else haploid), it is a T2T contig. As a contig, it is by definition gapless. If you have phased/haploid assembled sequence spanning from one telomere to the other, but it is interrupted by one or more gaps, it is a T2T scaffold. A T2T assembly contains T2T contigs, however as yet no consensus exists as to a minimum number/fraction of chromosomes recovered as T2T contigs to so-qualify an assembly.

Types of assembly/assembly output

For organisms with a single copy of their genetic code – e.g. prokaryotes and haploid eukaryotes, and for extraordinarily inbreed strains/cultivars used for research and agricultural, genome assembly outputs are straightforward: you (ideally) get one copy of the genome assembled which exactly corresponds to the genome sequence in each cell. For diploid or polyploid organisms, it becomes more complicated, with multiple types of assembly output possible. The original human reference genome was also a single copy of a human genome, made by arbitrarily choosing one allele at any polymorphic locus. Such an assembly is known as a collapsed assembly, and these assemblies are still broadly used as species or population references against which to map reads, call genetic variants, and represent these variants in a contig:position coordinate system.

With the advent of long reads and long read assemblers, a simple extension to the collapsed assemble was made for diploid organisms: the primary-alternate assembly. In these assemblies, over each polymorphic locus in the “primary” collapsed assembly, an “alternate” sequence is output, containing the allele found in the other haplotype. The alternate assembly is always significantly more fragmented than the primary assembly, however, because it breaks wherever a polymorphic region gives way to a homozygous region.

A logical extension to the primary-alternate assembly is the dual assembly, a.k.a. the partially phased diploid assembly. Essentially the only way this differs is that the homozygous sequence is copied over from the primary assembly to connect the alternate assembly sequences on either side, leading to two highly contiguous and complete copies of the genome assembled. Neither copy corresponds to chromosomal sequence actually present in your sample, however, as haplotypes are randomly shuffled on either side of each homozygous region (hence “partially phased”). Every single allele will be present in one or the other assembly, but each assembly will have a mix of paternal and maternal alleles within the same contigs.

We can use orthogonal data, including parental sequencing data or chromatin conformation capture sequencing (i.e. Hi-C or Pore-C ), to figure out which allele in a polymorphic region belongs to the same physical chromosome (i.e. haplotype) as an allele in another polymorphic region – indeed as all alleles in all polymorphic regions on the same contig. This process is known as assembly phasing, and it yields a phased or haplotype-resolved assembly. Long range phasing data can come from long reads (Pore-C or Oxford Nanopore parental data) or short reads (Hi-C or short read parental data). If parental sequencing data is used, we will have one assembly with just maternal chromosomes, and another with just paternal chromosomes. If chromatin conformation capture sequencing is used, we will have a mix of paternal and maternal chromosomes in each assembly – however for each chromosome the assembly will only have one or the other haplotype.


T2T_Know_how_figure_3_key_assembly_types

Figure 3. Illustration of the four key assembly types.

Assembly evaluation and quality assessment

Assembly quality has four separate, partially related axes: contiguity, completeness, structural accuracy, and base-level accuracy.

Contiguity is the measure of gappy-ness/brokenness: the longer the sequence between gaps/breaks, the higher the contiguity. This has typically been measured by contig N50: the maximum contig length on which you could have 50% of the genome; i.e. if you ordered all your contigs and added them up from longest to shortest, the length of the contig on which you would surpass 50% of the total. This is preferred over, say, mean or median contig length or number of contigs as all of these are highly sensitive to large numbers of small contigs at the tail end of the distribution. This metric worked very well when most chromosomes were recovered with many gaps, however now that we are recovering most chromosomes in one or two contigs, N50 is nearly maxed out and the remaining variation can be strongly driven sample characteristics (e.g. XX vs XY or random variation of chromosome length) or by which chromosomes the gaps fall on. As such, the number of T2T contigs can be a useful metric for comparing assembly contiguity and is our primary contiguity metric for T2T assemblies. Contiguity is sometimes also measured for scaffolds, with equivalent metrics (scaffold N50 and number of T2T scaffolds). As Hi-C or Pore-C data by itself in theory allows all chromosomes to be assembled to single scaffolds, scaffold level contiguity metrics may be less useful in the modern era.

Completeness is how much of the genome you have recovered, typically expressed as a % of the known sequence of your sample if your sample is well characterized, or else by a % of known single copy ortholog sequences (e.g. BUSCO genes – in which case the metric is called the BUSCO score) if a ground truth for your sample is not available. Very high contiguity usually means good completeness, and both metrics are usually essentially maxed out in a reasonably good Hifiasm (ONT) or Verkko assembly (the amount of sequence remaining in the few gaps tends to be negligible). However, this is a useful metric for comparing legacy assemblies or those produced with other sequencing technologies.

Structural accuracy is looking at whether the sequences were put together correctly overall, and can only be accurately measured by comparing back to a ground truth via whole genome alignment. The more complete, contiguous, and closely related the ground truth, the more precisely you’ll be able to measure structural accuracy. The HG002 human sample has the ultimate ground-truth: a manually curated fully T2T assembly with ca. Q70 base-level accuracy (see next section) confusingly named the Q100 T2T assembly (currently on v1.1). Using this assembly, we have never found a single translocation error (i.e. a completely wrong connection) in a single ultra-long read based Verkko or Hifiasm (ONT) assembly. Typically we find a few dozen large insertion/deletion errors that could be considered structural errors in these assemblies, although these tend to be in highly repetitive satellite DNA regions where even this near-perfect ground truth may be slightly suspect. In Hifiasm (ONT) assemblies from shorter reads (30-40 kb read length N50), we see up to a few hundred large insertion/deletion errors, the vast majority being in centromeric satellite units. We also typically see one or two translocation errors, always in rDNA arrays (i.e. always connecting the wrong acrocentric p-arm to an acrocentric q-arm). It is worth noting that to this date no one has ever correctly assembled a human rDNA array, and other software will not connect p-arms and q-arms at all.

Finally, we have base-level accuracy. This metric measures the total number of single base errors in an assembly as a (usually inverted) fraction of all the bases therein, i.e. how likely each base is to be correct. While it sounds straightforward, it is in fact probably the most tricky metric to properly evaluate. The ideal method is whole genome alignment, which we use internally for measuring base-level accuracy in the HG002 sample. However, you must have a sample-matched truth set or else your divergence will be dominated by genetic divergences rather than errors.

Change log

Date Version Changes made
July 2026 V2 Document updated to match T2T method offering overhaul
February 2025 V1 Document release
Oxford Nanopore Technologies, the Wheel icon, AmPORE-TB, EPI2ME, GridION, MinION, MinKNOW, PromethION, P2 Solo, and P2 are registered trademarks or the subject of trademark applications of Oxford Nanopore Technologies plc in various countries. Information contained herein may be protected by copyright, patents or patents pending of Oxford Nanopore Technologies plc. All other brands and names contained are the property of their respective owners. Oxford Nanopore Technologies products are RUO. Products labelled/branded as Oxford Nanopore Diagnostics may be RUO or may be regulated as in‐vitro diagnostic devices in some jurisdictions, please check individual product labelling. ONT plc is a member of the producer compliance scheme run by ERP UK Ltd, who manage the submission of documentation in support of WEEE compliance for ONT plc’s manufacture and supply of Electrical and Electronic equipment in the UK. ONT’s WEEE PRN is WEE/MM3828AA.

Last updated: 7/29/2026

Document options

Language:

Getting started

Buy a MinION starter pack Nanopore store Sequencing service providers Channel partners

Quick links

Intellectual property Cookie policy Corporate reporting Privacy policy Terms, conditions and policies Modern slavery policy Accessibility

About Oxford Nanopore

Contact us News Media resources & contacts Investor centre Careers BSI 27001 accreditationBSI 90001 accreditationBSI mark of trust
English flag