Wednesday, 18 July 2012

Illumina's Nextera capture: is this the killer app?

Illumina have released new Nextera Exome and Nextera Custom enrichment kits. These combine the rapid and simple sample prep of Nextera to in-solution genome capture and provide a straight-forward two-and-a-half day protocol. Add one day of sequencing on 2500 and you can expect your results in under one week.

I think many people did not expect exome costs to drop so fast. Ilumina's reduction of 80% was welcomed by many groups who wanted to run high numbers of exomes but could not when faced with the high costs and complex workflow without pre-capture pooling that other products had. Processing exomes is getting much easier on all fronts. Nimblegen compare their EZ-cap3 to Agilent in the EZ-cap v3 flyer, they did not include Illumina in thecomparison but they did use Illumina sequencing. It looks like Illumina can't lose whatever exome kit you buy!

A sequenced Nextera exome will cost £250 or $400 with 48 sample kits costing about £4000 and a PE75bp lane about £1000. That's cheaper than some companies are selling their exome or amplicon oligos for!

Nextera capture, how does it work: My lab beta-tested the new kits for exome and custom capture. The workflow was as simple as we expected and the data were high-quality. Nextera sample prep still uses just 50ng DNA and takes about three hours. The exome is the same 62MB but now with 12-plex pre-enrichment sample pooling, and still has two overnight captures.

We are still completing processing on our HiSeq of the full data set, but the initial analysis showed similar on and off-target results when compared to TruSeq, there was a higher duplication rate than we would normally see.

Below are three slides I put together for a recent presentation where I discussed our experiences with Nextera capture. They show that Nextera prep is simple (slide1), that QC can be confusing (slide2) but that results are good and costs are low (slide3). As low as £0.01 per exon!




Does Nextera custom capture kill amplicon-sequencing: You can design custom capture kits using Illumina's DesignStudio (reviewed here). The process is pretty straight-forward and a pool of oligos is soon on its way to you.

Illumina's Nextera capture only requires 50ng of DNA while TruSeq custom amplicon needs 250ng, so Nextera lets you capture more with less DNA. Small Nextera captures run at very high multiplexing on HiSeq 2500 are likely to be far cheaper than low-plexity amplicon screens on MiSeq. If Nextera custom capture can move to 96-plex pre-capture pooling then the workflow gets even easier (see the bottom of this thread for some numbers supporting this idea).

It will be interesting to see how the community repsinds to Nextera capture. If it takes off then Agilent, Nimblegen, Multiplicom, Fluidigm and others are going to be squeezed. There was a recent paper (MSA-cap seq) from the Institute of Cancer Research describing a low-input 50ng prep for Agilent capture, others like Rubicon and Fluidigm are aiming for the low input space, even single cells.

I am sure we have a lot more to look forward to for the second half of 2012.
Although I still don't see a $1 sample prep!


PS: One comment I have sent back to the development team is that the protocol requires the same 500ng library input into capture for exome and custom capture. This means the ratio of target to probe is much higher for custom capture and might be reduced. This would allow users to run custom capture with much less PCR. Alternatively if users stick with a 500ng input to capture then they might be able to increase pooling to much higher levels.

Exome capture uses 350,000 probes for 62Mb and 500ng capture input, giving 1.4ng per 1000 probes.
Custom capture uses 6,000 probes per 1Mb and 500ng capture input, giving 83ng per 1000 probes.
Can we run 96-plex pre-capture pooling?

Another thing I would like to see is the impact of completing only one round of capture. I'd assume we'd see a higher number of off-target reads. However saving a day when sequencing is so cheap for a small custom screening panel could be well worth it.

Wednesday, 11 July 2012

How NGS helped a physician scientist beat his own leukaemia

An amazing story was featured in the New York Times, the article follows Dr. Lukas Wartman from Wash U who is a leukaemia researcher who developed leukaemia himself.

The Genome Institute at Wash U made a concerted effort to find out what was behind Dr Wartman’s disease, performing tumour:normal and tumour RNA-seq analysis. He was fortunate to be included in a research study that was ongoing at Wash U, although that creates all sorts of ethical issues around who can access treatment and who cannot.

From the sequence analysis they found that FLT3 was more highly expressed than usual and could be driving his leukaemia. The drug Sutent had recently been approved for treating advanced kidney cancer, and it does it by inhibiting FLT3. Howwever it had never been used for leukaemia and unfortunately is costs $330 a day. Dr Wartman’s insurers and Pfizer turned him down for treatment.

The doctors he works with chipped in to buy a months supply and the treatment worked. Microscopic analysis showed the blood was clean, flow cytometry found no cancer cells and FISH was clear as well. Dr Wartman is back in remission and I certainly wish him all the best.

I'd love to hear back if circulating-tumor DNA analysis of his plasma is used as a monitoring test.

Tuesday, 3 July 2012

Is genomic analysis of single cells about to get a whole lot easier?

A couple of months ago Fluidigm disclosed the latest addition to their microfluidic chip technology, the C1 single-cell analysis system.

At the disclosure in May the details of the C1 were a little sketchy but the system was planned to take cell suspensions, separate and capture single-cells for nucleic-acid analysis on the Biomark system. The disclosure also hinted at the ability to process single-cells for NGS applications such as transcriptome and copy-number analysis. One rather worrying detail was the fact that the C1 was lumped together with the Biomark and FACS instruments as far as likely cost is concerned.

This means it is probably going to be an expensive instrument, which could limit its uptake. Like many sample prep systems only a few labs will have the capacity to run the box at full tilt. This often makes the investment hard to justify and labs end up sharing systems or using a service.
 

How does the system work: The C1 will take cells in suspension and isolate 96 single-cells. On the system you will be able to stain and image captured cells to determine what sort of cells are present from your population. Cells are then lysed and you can perform molecular biology on the nucleic acids (mRNA RT and pre-amp for now). Finally nucleic acids are recovered from the inlet wells for further analysis.



What is most interesting from my perspective is the ability to tie the C1 to NGS sample prep, either through transcriptome and/or Nextera based library prep, targeted resequencing through Access Array (see our recent paper for ideas) or WGA. If we can prepare libraries with 96 barcodes then I can see some wonderful experiments coming out of work in Flow Cytometry and Genomics labs.

What do I want to do with it: There are so many unanswered questions in biology that are hampered by using data from heterogeneous pools of cells; e.g. Tumours. Being able to dissociate cells, sort by flow cytometry and analyse with NGS is going to be powerful. From my experience of flow cytometry it is nearly always the case that any population of cells can be further subdivided. Being able to perform copy-number, mutation profiling or differential gene expression analysis on these populations is going to help get a better understanding of how much subdivision is required. Being able to capture perhaps 200 circulating tumour cells from a patient and sequence each of these is going to help us understand cancer evolution and metastasis. There are so many possibilities.

The C1 may also help one complex type of experiment that is often not replicated highly enough, low cell number gene expression. Some of my users provide flow sorted cells from populations that vary in number. Being able to run a pool of perhaps 100 cells from each population in a single C1 chip and get high quality GX data is going to remove much of the technical bias and enable more powerful statistical analysis.

Unfortunately this is going to mean lots of sequencing. Even if we can get away with 2M reads per sample for differential gene expression (genes, not exons or isoforms), then we’ll need to run a lane per C1 chip. It is not clear if we be able to generate copy-number analysis of sufficient quality from 0.25x coverage of a genome, although I have some ideas about that I'll post about later.

However for groups working on model organisms where cells can be grown in suspension or dissociated before analysis the C1 could be a phenomenal advance. Imagine individual C.elegans dissociated, cells identified by labelling, tagged by transcriptome library prep and sequenced. Fate-mapping on steroids anyone!

My lab has been working with Fludigm for two and a half years on the Biomark and Access Arrays and we did hear early on about their single-cell projects. I am excited to see a product is now available and look forward to working with the technology. Now all I have to do is find out just how much it will cost to complete an experiment!

Friday, 29 June 2012

Epigenome-seq state-of-the-art: but is hmC worth the effort for everyone?

A couple of recent papers have demonstrated the ability to distinguish between 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) using modified bisulfite sequencing protocols. These methods are likely to make a real impact on epigenomics when combined with Bis-seq. By sequencing two genomes, Bis-seq for mC and oxBS-seq or TAB-seq for hmC a fuller picture of methylation will emerge. However it is not clear how important biologically the additional data will be or how much it is worth to researchers.

Trying to complete this kind of experiment today on mammalian genomes requires quite a lot of sequencing muscle. Bis-seq depth guidelines are lacking but most people would aim for 50x or greater coverage. This suggests a 100 fold genome coverage per sample, or an expensive experiment. This could leave the approaches as a niche application for people with a strong focus on epigenetics.

oxBS-seq: The University of Cambridge’s method prevents hmC from being protected in a normal bisulfite conversion. 5-hmC is oxidised such that upon bisulfite treatment the bases are converted to uracils. See Quantitative sequencing of 5-methylcytosine and 5-hydroxymethylcytosine at single-base resolution.

TAB-seq: The University of Chicago’s method differs by protecting 5-hmC from conversion. 5-hmC is protected from oxidation by using beta glucosyltransferase whilst 5-mC’s are oxidised such that during bisulfite treatment they behave as if they were not methylated and are converted to uracils. See Base-resolution analysis of 5-hydroxymethylcytosine in the Mammalian genome.

How to make the methods accessible to all: An area that the two methods may have an impact on more quickly is for capture-based studies in cell lines where material is not limiting. It would be possible to create a PCR-free library and perform capture for an exome (or regions of equivalent genomic size) then follow this capture with bisulfite conversion and sequencing. This MethCap-seq (BisCap, oxBS-Cap or TAB-Cap) method could allow larger sample numbers to be run at the depth required without being too expensive.

We have been testing a Nextera based capture prep in my lab which potentially would allow this to be done on very small inputs, however the method is not yet released and requires quite a bit of amplification.

Another alternative would be to combine the methods with the recent Enhanced Reduced representation bisulfite sequencing (ERRBS) published by Maria E. Figueroa’s group in PLoS Genetics.

Wednesday, 27 June 2012

Has Qiagen bought the Aldi* or Wal-Mart* of DNA sequencing companies?

This update was added on 11 July 2012: An article on GenomeWeb made interesting reading as it talked about the current litigation between Columbia University and Illumina.  In it Thomas Theuringer (Qiagen PR Director) said "We support Columbia's litigation against Illumina, and are also entitled to royalties if we prevail". Maybe the royalties will be worth the reported $50M they paid?


Qiagen has just purchased Intelligent Bio-Systems and is, for now, the new kid on the NGS block. But what does IBS offer, what are the MaxSeq and Mini-20, can they compete against the monster that is Illumina and who owns the rights to SBS chemistry?

The Max and Mini instruments have been on my radar for a year or so now and I really did not think they would compete against Illumina when I first heard about them. I think one has been installed in Europe.

The instruments use technology from the noughties with maximum read lengths of 55bp and 1/3rd the yield of HiSeq. If the instrument and running costs are cheap enough users might be tempted. However a cheap alternative to a gold standard often turns out to be a disappointment.

If the title reads a little harsh then I’ll explain some of my thinking below.

*Aldi and Wal-Mart have a reputation in the UK for cheaper and lower quality products, they offer an alternative to supermarkets like Sainsburys. You should not make any NGS purchasing decisions based on a personal preference of where I like to do my weekly shop!

What did Intelligent BioSystems offer Qiagen and is there space in the NGS market for another player: Qiagen wants to get into the diagnostic sequencing space. However it is not clear if a Qiagen sequencer can compete against Roche, Life and Illumina in the research or clinical space. Clinically Roche already have a strong diagnostics division even if their 454 technology appears to be suffering with the strong competition of MiSeq and PGM. Illumina have thrown their weight behind clinical development and have a reputation for investing in R&D so expect them to become a major player. Life have a reputation for delivering products that work (lets ignore SOLiD) and the continued development of PGM and now Proton has got to keep them in a strong position.

Can Qiagen take any market share: The instrument will have to work robustly, deliver high quality data, be competitive on cost and include bioinformatics solutions. It is not clear what will differentiate the Qiagen instrument from others already widely adopted.

Although I think it is going to be tough for Qiagen it is not impossible; and if there is a $25B diagnostics market (UnitedHealth Group’s personalized medicine paper suggests that the molecular diagnostics market could grow to $25bn by 2021), even a 5% share of this is equates to $1.25B!

What will the Qiagen sequencers look like:
The sequencers offered by IBS make use of sequencing-by-synthesis technology licensed from Columbia university. This technology was published in 2006 by Jingyue Ju in Nicholas Turro’s chemical engineering group at Columbia University. Their PNAS paper describes a system that will be familiar to anyone using Illumina’s SBS. 3-O-allyl-dNTP-allyl fluorophores are incorporated by DNA polymerase, imaged and then cleaved with Palladium catalyzed deallylation ready for the next cycle. They sequenced 13bp in the 2006 paper at almost the same time as Illumina were buying Solexa for $600M.

Perhaps the most intriguing aspect of the technology is that the flowcells are reusable! Sounds great, but will clinical labs see this as a benefit over disposable consumables? I think not as there is too much risk a sample will become contaminated. The same goes for removing the need to barcoding on Mini-20 by using a 20 sample carousel. Barcoding is useful in labs just for sample tracking even if you end up doing a run in a single lane/flowcell.

Max-seq: In the brochure they claim that “thousands of genomes have been sequenced utilizing 2nd generation technologies, such as the MAX-Seq”, I would argue that 10,000’s of genomes have been sequenced but I am not aware of a single Max-seq genome to date.

Library prep uses ePCR bead-based or DNA nanoball (Rolony) methods. Libraries are loaded onto a flow-cell and sequenced with SBS chemistry. The instrument has dual-flowcells and each will generate 100M reads per lane in single or paired-end format at 35 or 55bp length, with >80% of bases being Q30 or higher.

The Azco website suggests that the Max-seq is thousands of dollars cheaper than a SOLiD or HiSeq instrument, and that run costs are 35% cheaper than Illumina or ABI. If you only get 25-50% of the data this make the system cost more like twice as much per run for lower quality data and much shorter reads.

Mini-20: I am not certain the Mini-20 actually exists yet. The brochure on the Azco website has a picture of the Max-seq with a line drawing of a flowcell carousel. The carousel should allow loading up to 20 flowcells and run up to 10 samples per day (SE35bp). A flow cell will generate 35M reads in single or paired-end format at 35 or 55bp length, with >80% of bases being Q30 or higher and 4Gb per flowcell.

Cost per run is predicted to be about $300 per flowcell. However it is not clear what the price would be if you wished to dispose of the reusable flowcells after a single use.

Those numbers do not add up to me but they must have made sense to Qiagen.

Friday, 22 June 2012

Improving small and miRNA NGS analysis or an introduction to HDsRNA-seq

Small RNA biases have been very well interrogated in a series of papers released in the last 12 months. The RNA ligation has been shown to be the major source of bias and the articles discussed in this post offer some simple fixes to current protocols which should allow even better detection, quantification and discovery in your experiments.

Small RNA plays an important regulatory role and this has been revealed by almost every method possible that can be used to measure RNA abundance; northern, real-time qPCR, microarrays and more recently next-generation sequencing. These methods do not agree particularly well with each other and the most likely candidate issue is technical biases of the different platforms.

Even though it has its own biases, small RNA sequencing appears to be the best method available for several reasons, it does not rely on probe design and hybridization, you can discriminate amongst members of the same microRNA family and you can detect, quantitate and discover in the same experiment (Linsen et al ref).

Improving small RNA sequencing: As NGS has been adopted for smallRNA analysis focus has appropriately been made on the biases in library preparation. Nearly all library prep methods use ligation of RNA adapters to the 3’ and 5’ ends of smallRNAs using T4 RNA ligases, before reverse transcription from 3’ adapters and amplification by PCR. However RNA ligase has strong sequence preferences and unless addressed these lead to bias in the final results of sequencing experiments.

All four of the papers below show major improvements to RNA-seq bias for small RNA protocols.

I particularly like the experiments performed in the Silence paper using a degenerate 21-mer RNA oligonucleotide. Briefly the theory is that a 21-mer degenerate oligo has trillions of possible sequence combinations and that in a standard sequencing run each sequence should appear no more than once as only a few million sequences are read. The results from a standard Illumina prep showed strong biases for some sequences that were significantly different from the expected Poisson distribution, and where almost 60,000 sequences were found more than 10 times instead of once as expected (the red line in figure A from their paper reproduced below). When they used adapters where four degenerate bases were added to the 5′ end of the 3′ adapter and to the 3′ end of the 5′ adapter, they achieved results much closer to those expected (blue line).


I don’t think we should be too worried about differential expression studies as long as the comparisons used the same methods for both groups, the results we have are probably true. However we may well have missed many smallRNAs because of the bias and our understanding of biology is likely to be enhanced by these improved protocols.


The recent papers:

Jayaprakash et al NAR 2011: showed that RNA ligases have “significant sequence specificity” and that “the profiles of small RNAs are strongly dependent on the adapters used for sample preparation”. They strongly suggest modifications to current protocols for smallRNA library prep using a mix of adapters, "the pooled-adapter strategy developed here provides a means to overcome issues of bias, and generate more accurate small RNA profiles."

Sun et al RNA 2011: "adaptor pooling could be an easy work-around solution to reveal the “true” small RNAome."

Zhuang et al NAR 2012: showed that the biases of T4 RNA ligases is not simply sequence preference but affected by structural features of RNAs and adapters. They suggested "using adapters with randomized regions results in higher ligation efficiency and reduced ligation bias".

Sorefan et al Silence 2012: demonstrate that secondary structure preferences of RNA ligase impact cloning and NGS library prep of small RNAs. They present “a high definition (HD) protocol that reduces the RNA ligase-dependent cloning bias” and suggest that “previous small RNA profiling experiments should be re-evaluated” as “new microRNAs are likely to be found, which were selected against by existing adapters”, a powerful if worrying argument.

Monday, 18 June 2012

Even easier box plots and pretty easy stats help uncover a three-fold increase in Illumina PhiX error rate!

One of the things I wanted to do with this blog was share things that make my job easier. One of the jobs I often have to do is communicate numbers quickly and effectively and a box plot can really help. I also have the same kind of troubles most people face with statistics, I find it hard! In this post I will discuss the GraphPad Prism package from GraphPad. This allows you to use stats confidently and make lovely plots (although annotating them is a nightmare). Recently the statisticians in our Bioinformatics core gave a short course in using GraphPad Prism.  I thought I'd explain the box plot in a little more detail and tell you a bit about GraphPad.

Previously I had shown how to create a box plot using Excel. I went down this route because I did not have time to learn a new package and Excel is available almost everywhere. However the result is less than perfect and it is hard work, indeed one major reason for writing the previous blog post was so I had somewhere to go next time I needed to create a plot! Statisticians like box plots as they can get across a lot more than just the mean and also say something about the size of the population being investigated.

Explaining box plots: A box plot is a graphical representation of some descriptive statistics; generally the mean and one of the following; standard deviation, standard error or Inter Quartile Range. A dot box plot is a version that allows these figures to be represented along with all the data points which allows the size of the population to be clearly seen. This helps enormously when comparing sample groups and deciding if a change in mean is statistically significant or not.

Dot box plots rule!

In figure 1 very similar mean and standard error are plotted with each dot representing a sample, see how the removal of some data does not significantly affect the "results" but seeing the data allows you to make a call on how much you are willing to trust it.


Figure 1

In figure 2 you can clearly see that there appear to be some "outliers" in group 2 but these have no affect on the results as the number of measurements is so high compared to group 1. Deciding if any "outliers" are in group 1 is much harder as the number of samples is so much lower. Removing outliers is really hard, and our statisticians generally advise against it.

Figure 2

GraphPad Prism: The software costs about $300 for a personal license. It might be a lot when budgets are tight but an academic license is not so expensive when shared across a department or institute. I’d certainly encourage you to take the plunge. It very quickly allows you to produce plots like the ones in this post as well as run standard statistical tests, and a whole lot more I won't go into. Take a look at their product tour if you want to find out more.

PhiX error rates: I used GraphPad to investigate an issue I had suspected for a while. We have been seeing a bias in the error rate on Illumina sequencing flowcells where lane one appeared to be higher than other lanes. Whilst the absolute numbers are not terrible and all lanes pass our QC there may be a real impact on results if this is not taken into account; mutation calling single lane samples and comparing tumour to normal for instance.

I took one months GAIIx data (8 flowcells) and plotted error rate for each lane. Entering the data into GraphPad is the most annoying bit and I usually copy and paste from Excel. However generation of the statistics and plots (figure 3) took about three minutes from start to finish.

A one way ANOVA with a Bonferroin correction showed how significant the differences were, with a very significant difference between lanes 1&2 and the rest. In fact there appears to be more of a gradient across a flowcell as lane 2 is affected, but at a lesser degree to lane 1.

A 2 way ANOVA allowed me to determine that in this data set lane accounted for 80% of the variance and instrument only 2.5%.

Figure 3
The biggest headache with GraphPad is the woefully inadequate annotation of graphs. Quite simply you will have to get an image out of the software and into Illustrator or PowerPoint. I guess if they are making the stats easy we should not complain too much.

I am using GraphPad on a weekly basis and for most reports where I have to summarise larger datasets. Why don't you give it a try.

PS: I'll let you know what Illumina say about the error rates. Please tell me if you've seen similar.

Wednesday, 30 May 2012

Don’t bother with a biopsy; just ask for a drop of blood instead.

I’d like to highlight some research that I have been involved in, which has just been published in Science Translational Medicine; Noninvasive identification and monitoring of cancer mutations by targeted deep sequencing of plasma DNA. The work makes extensive use of Fluidigm’s Access Array and Illumina sequencing, technologies that we have been running in my lab for over two and almost five years respectively. I’m proud of this work and I hope you like it as well.

Liquid biopsies for personalised cancer medicine:
Tumours leak DNA into blood and the SciTM paper shows how this circulating tumour DNA (ctDNA) can be used in cancer management as a “liquid biopsy”. The study by Forshew et al demonstrates the feasibility of testing and detecting mutations present at about 2% MAF (Mutant Allele Frequency), in multiple loci, directly from blood plasma in 88 patient samples. The method has been christened Tagged Amplicon sequencing or TAm-seq.

We know that specific mutations can impact treatment, e.g. ErbB2 amplification & Herceptin, and BRAF V600E & Vefurabenib, etc. Many cancers are heterogeneous with metastatic clones differing from each other and/or the primary, and biopsying all tumour locations is unrealistic for most patients. Understanding this heterogeneity in each patient will ultimately help guide personalised cancer medicine.

ctDNA however has not been easy to assay. It is usually fragmented to about 150bp and is present in only a few thousand copies per ml of blood. There has been a recent explosion in neonatal test development since the publications from the Lo and Quake groups. Whilst people were aware of ctDNA for many years it is only similar technological advances that allow us to assay it in such a comprehensive manner.

How does the liquid biopsy work: ctDNA is first extracted from between 0.8 and 2.2ml of blood using the QIAamp Circulating Nucleic Acid kit (Qiagen). Tailed locus-specific primers are used for PCR amplification and the loci targeted in the paper account for 38% of all point-mutations in COSMIC. 88 patient plasma, a couple of controls and 47 FFPE samples were tested in duplicate clearly demonstrating the utility and robustness of their method.

Each sample is pre-amplified in a multiplex PCR reaction to enrich for all targeted loci. This “pre-amp” is used as the template on a Fluidigm Access Array where each locus is individually amplified in 33nl PCR reactions that are recovered from the chip to produce a locus-amplified pool from each sample. A universal PCR adds barcodes and flowcell adapter sequences. 48 samples were pooled for each Illumina GAIIx sequencing run achieving 3200 fold coverage. We recently ran a library with 1536 samples in it from a collaborator using the technology in my lab. The potential of 12288 samples analysed in a single HiSeq run is astonishing.

How sensitive is the liquid biopsy: The paper presents results from a series of experiments to test the sensitivity, false discovery rate (using mixed samples) and validity (using digital Sanger-seq) of the TAm-seq method. ctDNA can be successfully amplified from as little as 0.8ml of plasma, far easier to get than a tissue biopsy! They were able to detect mutations at around 2% MAF with sensitivity and specificity >97%. The paper has some very good figures and I’d encourage you to read it and take a look at those.

Two of the figures stand out for me. The first shows results from one ovarian cancer patient’s follow-up blood samples in which they identified an EGFR mutation not present in the initial biopsy of the ovaries. When they reanalysed the samples collected during initial surgery they could find the EGFR mutation in omental biopsies. A TP53 mutation was found in all samples except white blood cells (normal) and acts as a control (Figure A from the paper).






They presented an analysis of tumour dynamics by tracking 10 mutations discovered from whole tumour genome sequencing, using TAm-seq in the plasma of a single breast cancer patient over a 16 month period of treatment and relapse (Figure D). They also demonstrated of the utility of TAm-seq by comparison to the current ovarian cancer biomarker, CA125 (figures B & C not reproduced here).

What does this mean for clinical cancer genomes: There are many reports of whole genome sequencing leading to clinically interesting findings and some labs have started formal exome and even whole genome sequencing on patients. Whilst there is little doubt that tumour sequencing is likely to be useful in most solid tumours it is still hard to see how this will trickle down to the 100,000 or so new cancer patients we see in the UK. Challenges include cost, bioinformatics, incidental findings and ethics.

The TAm-seq method has fewer of these challenges (although it is not good for detecting amplifications) and I think it is a really big step in clinical cancer genomics and hope it translates quickly into the clinic to inform treatment and prognosis. Perhaps it will be the first technique to make a big splash in personalised cancer genomics?

Hopefully this liquid biopsy will be quickly translated into the clinic.

Perhaps a future version might even turn up in your doctor’s surgery on a MinION in a few years time as a basic screening tool?

Saturday, 26 May 2012

The exome is only a small portion of the genome

Ricki Lewis (Genetic linkage blog) wrote a great guest post for Scientific American on what exome sequencing can’t do. It seems timely considering the explosion of interest in exome sequencing and exome arrays. Not so long ago most people I knew still talked about junk DNA, exome sequencing and exome arrays essentially allows users to ignore the junk to get on with real science. As Ricki points out exome analysis is a phenomenally useful tool but users need to understand what they can’t do to get the most from their studies.

Ricki listed 10 things exomes are not so good for, my list is a lot shorter at just 4.
  1. Regulatory sequence is missing (although this is being added, e.g. Illumina).
  2. Not all exons are included.
  3. Structural variants (CNV, InDel, Inv, Trans, etc) are not easily assayed with current exome products.
  4. No two exome products are the same.

Exome analysis has had a real impact, especially on Mendelian diseases that remain undiagnosed. However users need to remember they are only looking at a very small portion of the genome. Ricki puts it this way “the exome, including only exons, is to the genome what a Wikipedia entry about a book is to the actual book”.

I posted a month or so ago about choosing between exome-chip and exome-seq. The explosion in exome-chips has been an even bigger surprise then exome-seq. Illumina admitted that they had been overwhelmed with demand for their array products. It appeared to be pretty clear that exome-seq would take off as soon as the cost came down to something reasonable. However according to Illumina over 1M samples have now been run on exome chips!

Of course analysis of an exome is allowing studies to happen that would never get off the ground if whole genome sequencing were the only option. The cost and relative ease of analysis makes the technology accessible to almost anyone. As the methods and content improve over the next coupe of years this is going to get even easier.

The simplest thing for users to remember is that they are restricting analysis to a subset of the genome. This means that just because you don’t find a variant does not mean one is not lurking outside the exome; absence of evidence is not evidence of absence as statisticians would put it.

It is also helpful to remember that not all exomes are created equal. Commercial products are designed with a price and user in mind. Academic input is usually limited to a few groups and there are always other bits that could be added in. Illumina have done a great job including some of the regulome in their product but the commercial products are in a similar arms race to the one faced by microarray vendors a decade ago. Just because a product targets a bigger exome does not mean it is better for your study.

Exomes are well and truly here to stay. We'll probably see an exome journal soon enough as there is so much interest.

Thursday, 24 May 2012

AmpliSeq 2

Ion Torrent released AmpliSeq 2.0 a little while ago. The biggest change is an increase to 3072 amplicons per pool. I saw a Life Tech slide deck which had up to 6144 amplicons pre pool so there looks like more room for improvment.

Who knows when we might see an multiplex PCR exome!

PS: see a previous post about AmpliSeq for more general details.