Wednesday, 15 April 2015

Should you buy a NeoPrep (or any other NGS automation)

Illumina launched NeoPrep at AGBT and Keith Robison at OmicsOmics wrote a detailed summary of what Illumina say the NeoPrep is capable of (he compared this movie of the electro-wetting technology to video games of his, and my, youth), I thought I'd write down my thoughts on how this instrument might fit into labs like mine and possibly yours too.



The main selling point of NeoPrep is that it provides a one-stop solution for NGS library prep. The price point is pretty good (around £30-35k), and speed and quality  looked great in the data presented by Illumina at AGBT (by Gary Schroth and Kevin Meldrum, Illumina have run over 5000 libraries so far); so is NeoPrep a good option for every NGS lab? There are a couple of limitations I'll return to in a bit - but broadly speaking I can see that this system really could be a good, even a sensible, fit for many many labs running NGS. In a core lab or heavy NGS research lab the NeoPrep looks like  it will take some of the worry out of NGS library prep. And even is a lab that only does a dozen or so NGS experiments per year, removing the worry that library prep will go wrong with precious samples might be enough to warrant a purchase.

Because NeoPrep can be run in such a hands-off way, and because the quality and reproducibility of data are reportedly high (although we'll have to wait for  user reports over the next six months for confirmation of this), then a lab that is spending PostDoc or PhD time making small to medium sized numbers of libraries might well buy a NeoPrep where they would never have considered purchasing a liquid handling robot due to their complexity. This means Illumina might have hit the nail squarely on the head on this instrument.

For Core Labs the ability to offer library preps with an Illumina guarantee of quality is likely to be a positive, and something customers might approve of. And as Illumina release more library prep methods for NeoPrep the instrument might be perfect for those things you rarely get asked to do.

The positives: If the hands on time claims of just 30 minutes for sequencing ready libraries are true then NeoPrep is going to save people time and that is probably our most precious commodity. Labs like mine like big projects and we tend to batch smaller ones together, this can mean a wait for users while other samples come in to fill a 96well plate. NeoPrep would allow us to run projects as small as 16 samples. Alternatively it would allow us to provide automation solutions directly to users, running NeoPrep as a bookable instrument in the core.

The price per library is attractive and Illumina are aiming for parity with manual kits (which could generate a "why bother" attitude to manual prep). The TruSeq Nano kit, comes in at around $30 per sample, and TruSeq stranded mRNA around $55. GenomeWeb quoted Illumina as saying "the "fully loaded cost per sample" using NeoPrep would be around $75 per sample, which includes the cost of amortization of the instrument...assuming 1,000 samples are run per year." Compare that to the real cost of a top-flight post-doc spending two weeks making libraries for one experiment!

Price is not everything, so also encouraging is Illumina's data on reproducibility which appears to be very good in the results presented so far; concordance of 0.97 between RNA-seq libraries of varying inputs (100ng vs 10ng, and even 100ng vs 2ng). And the NeoPrep is being touted as requiring less starting material, in fact Illumina recommends starting with 25ng to 75ng of DNA for TruSeq Nano which is several thousand genome equivalents. For TruSeq mRNA, they recommend an input of just 25ng. They also tested down to 2ng but reported some drawbacks, including a lower yield and increased numbers of duplicates. I do worry about biases in low-input experiments and we previously showed a drop in sensitivity, although not specificity, as RNA input dropped from 100ng to 10ng in Illumina HT12 arrays. I guess we should repeat this experiment with TruSeq RNA on NeoPrep!

The negatives: The most obvious problem is that you can only run Illumina reagents on the NeoPrep and whilst they are competitively priced some competitors are significantly cheaper. The range of library preps is also very limited with just DNA Nano and TruSeq stranded mRNA at the moment (GenomeWeb quoted Illumina as saying the "PCR Free kit, followed by a "steady stream of protocols" [would begin] in the second half of the year and including targeted resequencing panels"). I know I'd like to see Nextera exomes ASAP; and knowing a timescale for ChIP-seq and ribozero would be great.

There is a huge range of methods that have been developed to run on Illumina sequencer (download Jaques Retief's amazing poster with almost 150 apps), most probably only a handful of these will ever see the light of day on a NeoPrep. But many of them use steps in Illumina's core library prep technology (end-repair, adapter-ligation), so if NeoPrep could be configured by users we might be able to make it do what we want. As an example we've been making some RNaseH libraries with NEB kits for ribo-depletion, we did this by eluting the ribodepleted RNA from Agencourt RNAClean XP beads directly into the 19.5 μl of Fragment, Prime, Finish Mix in the Illumina TruSeq stranded mRNA kit, then we carried on with the protocol from "Incubate RFP" to make multiplexed RNA-seq libraries.

If it is unclear what else might come, and when, then making a purchase decision is tougher. Perhaps more important is understanding what methods Illumina can not migrate to NeoPrep, I suspect anything that needs a gel is going to be difficult.

Some of the library prep kits from other companies are way better than Illumina's for specific applications. This is not because Illumina can't make those kits (unless IP stops them), but is probably more a case of Illumina looking to see which markets are largest, and possibly which competitors they want to crush. If someone does something Illumina can't, or won't, do then we'll not be running that on NeoPrep.

My biggest concern with NeoPrep is that it is a black (and white) box so user's don't need to know what is going on under-the-hood; you could equally say my labs library prep services. I am a strong believer that you (we) need to understand the library prep technology to innovate and NeoPrep might reduce the likelihood of new users doing their homework. It could also be argued that as long as you can ask a sensible question and can validate and interpret the results, then using a black-box to go from library-prep (NeoPrep) to results (BaseSpace or similar) does not matter. However much of the innovation in NGS methods has come from tweaks to library prep methods, and I'd hate to see a slow down in this space.

Lastly 16 samples at a time could be limiting if it is not easy to run a 96plex experiment in four runs! High-plexity sequencing rocks!

How to choose if NeoPrep is for you: Illumina have data on their blog and website, they also have an unbiased buyers guide to laboratory automation that is probably worth reading if you're thinking "should I buy a NeoPrep or a Agilent/Beckman/Hamilton etc?". Ultimately you need to consider the number and type of libraries you want to make over the next couple of years. You might decide that kits like Thruplex, KAPA Hyper+ or Lexogen make library prep so simple that you don't need a robot at all.

Your needs will ultimately guide you and my thoughts in this post are very squarely for Human genomics and transcriptomics. If you are working with small genomes then other technologies to reduce costs by orders of magnitude are probably more interesting e.g. high-plexity mRNA-seq with a single library prep.

Data from my lab: I'd love to be able to share RNA-seq data from my lab with you but I can't; because there appear to be no demo units available! However I did run through the demo instrument at AGBT and the run setup wizard was a longer version of the same wizards on HiSeq, MiSeq and CBot. I could see myself starting 16 RNA-seq libraries first thing in the morning and then doing the day-job while NeoPrep makes the libraries. I could also see a lab like mine wanting three instruments so we can run 48 samples in a batch, hopefully they will rack nicely to save bench space.

Wednesday, 8 April 2015

Rise of the Nanopore

Nature Methods recently carried a News & Views article from Nick Loman and Mick Watson: “Successful test launch for nanopore sequencing”, in which they discuss the early reports of MinION usage; including a paper in the same issue of Nature Methods (Jain et al). They recall the initial “launch” by Clive Brown at AGBT 2012; this caused a huge amount of excitement, which has been tempered by the slightly longer wait than many were hoping for. Nick ‘n’ Mick suggest that a “new branch of bioinformatics” is coming dedicated to the nanopore data (k-mers), which is very different compared to Sanger or NGS data (bases).

Jain et al: The paper in Nature Metohds from Mark Akeson’s lab at UCSC presents the sequencing of the 7.2kb genome of M13mp18 (42%GC) and reported 99% of 2D MinION reads (the highest quality reads) mapping to the reference at 85% raw accuracy. They presented a SNP detection tool that increased the SNP-call accuracy up to 99%. To achieve this they modelled the error rate in this small genome at high-coverage; 100% accuracy might be impossible in homopolymer regions where the transition between k-mers is very very difficult to interpret, but for much of the genome MinION looks like it will be usable. Whether this approach will work for targeted sequencing of Human genomes will be something I’ll be working on myself.

In the paper they also reported very long-read sequencing of a putative 50-kb assembly gap on human Xq24 containing a 4,861-bp tandem repeat cluster. They sequenced a BAC clone and obtained 9 2D reads that spanned the gap allowing them to determine the presence of 8 repeats, confirmed by PFGE and “short” 10kb-reads from fragmented BAC DNA. 

The future looks bright for MinION: Jain et al discuss the rapid rate of improvement in MinION data quality, and Nick ‘n’ Mick also mention this when talking about why they're so upbeat about the MinION (hear it directly, both are speaking at the ONT "London Calling" conference). Their main reason is the success of the MinION Access Program in its first year (e.g. Jain et al reported the increase in 2D reads due to changes in sequencing chemistry from June (66%), July (70%), October (78%) and November (85%); and Loman published a bacterial genome after just 3 months in the MAP demonstrating the improvements in chemistry); they also point out that very long-reads allow access to regions of the genome off-limits to short-read technologies; and they mention the hope of direct base-modification analysis, direct RNA-seq and protein sequencing. Jain et al also discuss the possibilities of detecting epigenetic modifications, etc. These all seem a very long way off to me, but with so many labs participating in the MAP who knows how soon we’ll be reading about these applications? 

Jain et al and Nick ‘n’ Mick both mention the miniature size of the MinION and its portability. It is certainly small, I accidentally took mine home after a meeting because it was in my pocket! If this portability can move sequencing from the bench-to-bedside then MinION could be the first point-of-care diagnostic sequencer. It may be premature to suggest this, but many cancer researchers would love to sequence DNA directly from blood with as little time in-between collection and sequencing – if Clive’s AGBT 2012 claim “that sequencing can be accomplished directly from blood” proves to be accurate then this may just be a matter of technology (mainly sample prep) maturation.

I agree that the future looks bright for MinION. ONT tried something quite different with the MAP, this was a risk but is one that seems to be paying off. Year two is likely to see many many more publications from the large number of MAPpers.

Disclosure: I am a participant in the MinION Access Program.

PS: You can find a few MinION’s on the Google Map of NGS.

Thursday, 2 April 2015

BGI sells off Illumina HiSeq instruments

Looks like I missed out on a bargain, the BGI has sold all their HiSeq 2500 instruments  on eBay, but the auction ended yesterday.


PS: Is the 454 on eBay for £7000 an April fool?

Tuesday, 31 March 2015

Book Review: Bang wongs "Visual Strategies for Biological Data"

I've written before about how much I liked Nature Methods “Points of View” by Bang Wong and I created a public Mendeley group so you could access the papers. I'd also said that having the articles collected together as a hard-copy version would be great.

Now available is Nature Collections: Visual Strategies for Biological Data "this
e-book collects the Points of View columns published in Nature Methods through February 2015, providing practical advice on effective strategies for visualising biological data to researchers in the biological sciences."
 
Enjoy. 






Friday, 20 March 2015

Oxford Nanopore MinION for ctDNA sequencing


A great poster at AGBT was presented by Boreal Genomics and available on the Nanopore wiki for MAPpers. In A nanopore liquid biopsy Patrick Davies describes their combination of the Boreal On-Target with ONT MinION sequencing to detect mutant allele fractions in ctDNA of sub 0.1%. I spoke briefly to Andre Marziali (Boreal Founder & CSO) about the work and summarise the poster here.

Saturday, 14 March 2015

Fancy working in my lab?


I've currently got three positions open in my lab and thought I'd use this blog as another way to get the message out to prospective candidates. Two people recently moved onto new jobs; one in Inivata (the first spin-out from CRUK-CI) and one at AbCam, and another person was recently promoted. We're also busy so we're also recruiting for a six-month temporary contract to help out with the sequencing services.

If you want to see what the lab does please take a look at our lab website and you may have seen us on Twitter.

The posts:

The posts will all be involved in providing Next-Generation Sequencing and library preparation services; including nucleic acid and library quant (with KAPA); setting up, monitoring and troubleshooting Illumina HiSeq, NextSeq and MiSeq sequencers; and library prep using a diversity of methods, such as Exome-seq, ChIPseq, and RNASeq - we do a lot of RNA-seq and Exomes. The senior post wil be responsible for the day-today operational management of the NGS service, and will work alongside their counterpart running the library prep services.

The Genomics core has been operational for 8 years and we've focused on NGS for 7 of those; it is an experienced lab doing exciting work with a diverse set of users from across the Cambridge, although our primary focus is on Cancer Research methods for scientists at the CRUK funded Cambridge Institute.

Please follow the links for details on applications rather than contacting me directly.

Thanks.

James.

PS: Closing date for all posts is 27 March 2015.


Friday, 13 March 2015

A better way to sequence exomes?

I caught up with a new company on the target capture scene, Directed Genomics, at AGBT. Their approach is based on a simple idea: if you want to sequence exomes, why not capture only exons?

Most exome-seq methods (Illumina, Agilent, Nimblegen) use oligo-baits to pull-down adapter-ligated fragment libraries, with fragments of 200-300bp. As exons are only 170bp long (80–85% Human exons less than 200bp Zhu et al & Sakharkar et al) we sequence lots of near- or off-target bases. These can be used (cnvOffSeq for instance), but are to some degree wasted sequencing.

The Directed Genomics approach: similar to other exome capture companies Directed Genomics also uses a probe hybridisation to targeted regions and/or exons, but applies this in a very different manner than we’re used to with standard exome capture. Two methods are presented in their recent posters; the first uses two probes, one at each end of the exon; the second uses a single probe hyb and random 5’end to create molecularly identifiable libraries. Current plans appear to be for custom panels, but hopefully they'll to build out to a whole exome panel over time.

Directed Genomics workflows

1: In their dual-probe method a short 50bp biotinyated-oligo probe is hybridised to fragmented gDNA at the 3’ end of an exon, the sequence upstream of this is then enzymatically digested and the 3’ hairpin adapter ligated. Next a second 50bp probe is hybridised to the 5’ end of the exon, the 5’ end is blunted and a 5’ adapter is ligated. Rather cleverly the hairpin adaptor ligated at the 3' end of the target links the target to the probe, allowing for a heat step in the second probe hybridisation without losing the target. Finally the 3’ hairpin is cleaved releasing products for PCR amplification and sequencing that contain only targeted exonic sequences. On-target rates of 97% were reported in their AGBT poster.

2: In their single-probe method a short 50bp probe is hybridised to fragmented gDNA at the 3’ end of an exon, the sequence upstream of this is then enzymatically digested and the 3’ adapter ligated. The probes is then extended to create the complementary strand and a 5’ adapter is ligated to the blunt end. This creates a library with random 5’ ends enabling a duplicate filtering step, unlike PCR approaches.

The protocols are both same-day 6-8 hours with around 1.5 hours hands-on time (according to the posters). Both allow a certain amount of, or all of the off-target sequence to be removed, reducing the amount of sequencing wasted. However the variation in exon length means that some sequence is inevitably lost.

Molecular IDs in cell free DNA: Their single-probe method creates libraries with in-built molecular ID. The random nature of the 5’ end should allow removal of all PCR duplication, without affecting biological duplication too much. Adding a  molecular identifier to the 3’ probe would increase this even further; and also bring molecular ID to the of the dual-probe method.

These molecular ID’s are likely to become increasingly important in methods to call low-frequency mutations in cell-free DNA applications, particularly ctDNA. Current methods make use of deep-sequencing to call mutations just below 1% MAF (mutant allele freq). However simply sequencing deeper may not be enough to get under 0.1%. A MAF of 0.1% would require sequencing to >10,000x to have enough mutant allele reads; and PCR, clustering and sequencing errors all make the detection harder.

Adding a molecular identifier should allow us to develop better statistical methods to call lower and lower MAF. Ultimately we aim to get to a point where we are restricted more by the presence of mutant alleles in a sample than by the technology used to capture and sequence them.

Directed Genomics and cell free DNA: The AGBT poster contained results from the Horizon Diagnostics Multiplex Reference Standard (link). Correlations of observed vs expected allele frequencies were >0.91. This is one of the first methods that can target mutant alleles with a single oligo, as compared to the two used for PCR amplicon sequencing, e.g. TAM-seq. It should mean an increase in sensitivity as more ctDNA molecules can be captured and amplified.

Directed Genomics expects to be launching later in 2015.

Thursday, 12 March 2015

Combining high-throughput CRISPR with in silico cancer drug development

In my last post I wrote about a computational screen of TCGA data and its use in repurposing approved drugs and/or finding new drug candidates for cancer patients. The work demonstrated the possibilities for finding novel treatments, but I also pointed to a cautionary Vemurafenib study that showed poor performance repurposing the drug in Colorectal cancer. As it becomes easier to identify novel therapies in a high-throughput manner, we need to develop methods to test these the are equally high-throughput. CRISPR knock-out or mutation of cancer drivers in multiple cancer cell lines or in tumour xeongrafts is one possibility - but most groups have carried out only a handful of knock-out or genome editing experiments.

Tuesday, 10 March 2015

In silico prescription of cancer drugs is likely to benefit patients - and can only get better

A fantastic paper just out in Cancer Cell: In Silico Prescription of Anticancer Drugs to Cohorts of 28 Tumor Types Reveals Targeting Opportunities from Nuria Lopez-Bigas's BioMedical Genomics lab in Barcelona.

They have developed an in silico prescription strategy by identifying the driver events in TCGA data, collating data on therapeutic drugs that target driver genes, and connecting  patients with driver mutations to potential therapies (see figure 1 from their paper below).
  • 40% (1635) of patients benefit from in silico prescription and repurposing of FDA-approved drug
  • 33.1% (1346) of additional patients benefit from in silico prescription and repurposing of drugs currently in clinical trials
  • 39% of patients could benefit from novel combination therapies
Figure 1 from Rubio-Perez et al (Cancer Cell 2015)
To identifying the driving events they took data from 28 cancers studied as part of TCGA and analysed all somatic SNVs, InDels, CNVs, fusions and RNA-seq differential gene expression. They found well over 400 genes that drive tumorigenesis via mutations, CNAs or gene fusions (data available at IntOGen http://www.intogen.org; Gonzalez-Perez et al., 2013a). Many of these driver events are in loss-of-function mutations that could be druggable, are present in many samples, and are not in well-established cancer genes. 25 driver events occur in at least 5% of tumours of at least one cancer type.

Understanding cancer biology is vital: Whilst exciting the results presented need to be taken with a pinch of salt, and one worries about the headlines journalists might be using in non-scientific media! 

ICGC/TCGA and other NGS-based cancer projects have discovered new insights into cancer biology. However the results need very careful evaluation (and clinical trials) before their impact can be stated, and, at least in the case of Vemurafenib, targeted therapy can fail when applied to a different cancer. Vemurafenib increases survival in up to 80% of melanoma patients with the BRAF V600E mutation (although many patients develop resistance). But when prescribed to BRAF V600E positive colorectal cancer patients only 5% responded (see Prahallad et al in Nature 2012). This was reported as being due to activation of EGFR, by inhibition of BRAF V600E, driving continued cell proliferation. EGFR is expressed at low-levels in melanoma so the feedback activation is not significant. examples like this, and we can expect more to be reported, demonstrate the need to understand cancer biology with respect to targeted therapeutics.

The future for this kind of analysis looks bright: ICGC/TCGA data sets are getting larger and richer, analysis algorithms continue to improve, data on basket trials is starting to be reported, companies like Foundation Medicine are developing tests to report this kind of result. Hopefully this kind of analysis will be routine in just a few years time.

Friday, 6 March 2015

10X Genomics: what's the fuss over phasing

At AGBT 2015 the big splash was clearly 10X Genomics and their new technology the GemCode "toaster"; presumably so called because of its diminutive size, and not because your microtitre plate is launched out the top nice and warm! The system is available to order now costing $75K, with a $500 per sample price. Using an input of just 1ng means users can test this even with precious clinical samples. Hopefully the improved structural variant detection 10X are promising will have a significant impact on cancer research, perhaps making translocation   discovery easier.