Wednesday, 27 March 2013

Finding your next post-doc lab is like getting a date for Saturday night (sort of)

A colleague and I were discussing the woes of post-docs trying to find their next job, or a tenure-track position. My colleague was relating the story of a mutual friend who recently published an excellent first-tier publication and was still having problems finding a junior group leader position. She’d said to me that “If so-and-so could not find a job what was the chance for the rest of us mere mortals” implying that you can only get on with multiple Nature, Science or Cell papers!

Tuesday, 26 March 2013

Imputation of LOH from 1x genome sequence

I would like to move from microarray based genotyping to NGS for copy-number and loss-of-heterozygosity analysis. The copy-number bit should be relatively simple and I hope we’ll have done this by the end of the year, however the LOH analysis is more complex as it requires high quality genotype calls and to get these from sequencing you need 10x coverage. Or do you?

Monday, 25 March 2013

Making NGS greener

Does your lab look like this? We get most of our deliveries on dry-ice shipped from European distribution centres. All the polystyrene and dry-ice are the tip of our energy consumption iceberg. Genomics sciences have as much environmental impact as just about everything else, and perhaps quite a bit more than other disciplines.

Tuesday, 19 March 2013

Money laundering 101 for core facility managers

It's that time of year again. The end of the financial year (for many of us) when most core facility managers get frantic calls asking if work can be billed now but done later. Lets face it this is simply a money laundering exercise but two things irk me when it comes round to this time of year.

Firstly why can't grant agencies simply be more flexible with how money is spent on a grant. If the time lines change and money is left over in one year there should be a reasonable degree of flexibility in moving it to another financial year.

Secondly, why can't PI's keep better track of their money. In some institutions budgets can run into many $100,000s or even $millions. Finding out in the middle of February that you have nearly $100,000 left that needs to be "spent" by March 31st does not give you much time to spend it wisely.

My two most frequently requested money laundering methods are:
  1. Buying reagent kits for use later in the year: a simple one to do as most companies will happily expedite reagents to get a large order. But please don't expect your core manager to spend a week negotiating over price, and do remember reagents have a shelf-life!
  2. Buying services ahead of time: again a simple one to do and remember that the service you asked for can also go out of date. What happens if the service provider stops running your favourite array? What happens if a technology changes and you're committed to what now looks like an exorbitant price?
What damage does the system do to science? From my perspective the biggest problem with the sometimes "use it or lose it" accounting is that we can make rushed, and at worst bad decisions. There is nothing like a tight deadline to force people into taking shortcuts. Perhaps only getting one quote instead of three, or not considering the cost of future commitments.

I was never trained as an accountant but for the past fifteen years I have been responsible for pretty large budgets. Most core managers and group leaders that I know are spending considerable sums of often public money and this last minute pressure is one I am sure we could all do without. A bit of extra financial training wouldn't go amiss for most of us either.

If you are laundering money for someone, or requesting monies to be spent at short notice do think about the impact. I have known one group that were faced with buying a new set of reagents because their early purchase went out of date. The samples were irreplaceable and my recommendation was to buy new kits. It was a painful decision and one we would not have needed to take had we been able to reserve the money for closer to when the samples were actually ready!

Wednesday, 6 March 2013

Derek and Clive live (sort of): Nick Loman chats to Clive Brown at ONT

Nick Loman had a nice cosy chat with Clive Brown at AGBT and tells us all about it on his blog Pathgogens: Genes and Genomes. It reads to me like a very reasonable assessment of why we have heard so little from Oxford Nanopore in the last 12 months. ONT have a very disruptive technology that could be months, years or decades from release, but when it comes it could knock the socks off Illumina and others.

I completely agree with Nicks comments that when ONT deliver the community will be quick to forget the lack of information. They hype is created as much by our own discussions as ONT's announcement followed by silence.

I was surprised to hear Clive say the ONT were caught off-guard with the interest in their AGBT 2012 presentation. No-one in the world could stand up at the biggest sequencing technology meeting their is, tell everyone about MinION and then expect us to nod serenely and go back to MiSeq vs PGM (the battle is clearly being won). We were excited, enthralled and desperate to get our hands on the technology first. No wonder we pesterd and griped, most of us probably felt like we might be missing out not to be in the early-access program (unless you're all in it and I'm the only one who isn't).

I, like Nick, sure hope I'm not on Clive's naughty list!

Although Nick may find himself on it not for his negative comments on ONT's strategy but more for his "I spotted a distinctive bald head" comment!

PS: Derek and Clive Live had a lot more swearing!

The joy of grant funding for NGS

At regular intervals throughput the year I am asked for some help with costing next-gen sequencing experiments in grant proposals. The request usually comes towards the middle of the week leaving only a day or two before the deadline but that's just how we work in academia!

One thing that has struck me over the last five or six years is how much more sequencing a grantee can do when the samples are ready compared to when the grant was written as there is usually a delay of months (even years). If you take a look at the graphs below showing how costs have dropped between Jan '10 to Sept '12 you hopefully get the idea.
Wetterstrand KA. DNA Sequencing Costs: Data from the NHGRI Genome Sequencing Program (GSP) Available at: www.genome.gov/sequencingcosts. Accessed [March 6th 2013].

Making research $$$ go further: A grant awarded in 2010 for $500,000 might be spent over 3-5 years. If the bulk of the grant is for sequencing genomes, and the grantee spends some time collecting samples the affect on the number of genomes the grantee can sequence can be huge. The graph below shows that 11 genomes could be sequenced for $500,000 in January 2010, wait six months and this rose to 16 genomes. Wait a year and you could have sequenced 24, two years 65 and if you could hold out for three years you could sequence almost 10x as many genomes for the same $500,000.


Will the NIH, Wellcome and others wake up to this: I've not looked at this before and it is rare to go through such a precipitous drop in the cost of a method. However most pundits are predicting the price to continue to drop and Illumina already announced big increases to capacity. Maybe we should be awarding the sequencing portion of a grant at 50% of what was asked for?

Tuesday, 5 March 2013

What other medical tests could you buy with a $1000 genome

There is lots of discussion about the $1000 genome, what it means, what its impact might be, etc.

First off though deciding what a $1000 genome means seems to be almost impossible given everybody uses different ways to calculate costs. In a soon to be published review article we pointed to the 2011 Sboner et al paper that stated a cost of $6500, and we suggest that today a genome costs about $4000 for 30x coverage. Whether you include instrument amortisation as well as data analysis costs makes a huge difference to the final figure. Is it really closer to $25,000?

Even so a "$1000 genome" can be compared to current medical tests and in this post I've put together a list of ten common tests ranging in cost as comparators. I was surprised at the costs of some tests and can see even NGS based tests coming in at well under $100 (about the same as an ultrasound), and perhaps even $10 in the future (at the same as a pregnancy test).

The price of medical tests:
 Robotic surgery to remove prostate £10,000
Standard course of chemotherapy £5000-10,000
BRCA test or Oncotype DX £2,800
Cancer gene sequencing (all exons) £1000
Overnight stay in Hospital £400
CT scan £400
MRI scan £300
Endoscopic biopsy£350
Muscle Biopsy £200
Ultrasound £100
Skin biopsy £50
X-ray (complex) £25
Full biochemistry profile £20
Full microbiology profile £15
Pregnancy test £7.50
Glucose test £5

Will consumer genomics make NGS even cheaper: If people respond to genomics and ask for it from their healthcare providers then costs could fall. A multiplex PCR that amplifies several cancer genes could be very cheap to run if the numbers are high and restricted to genes with relatively simple interpretations.

Can we do cancer NGS testing for as little as a pregnancy test?

PS: the costs of these tests come from a variety of sources but all are likely to suffer the same problem of how those costs were derived. Hopefully this is more Bramleys to Cox's rather than apples to oranges.

Monday, 4 March 2013

Oxford Nanopore chip announced!

Almost...

I had a very enjoyable trip to the Science Museum in London this weekend and whilst there was amazed to see an Oxford Nanopore DNA sequencing chip on display.

The word on the street after ONT's AGBT 2012 extravaganza has been more like a quiet grumble about progress, the hype has certainly not been lived up to, now that quiet grumble is beginning to get noisier! Everyone was hoping for their MinION USB sequencer and the Star-Trekesque possibility of truly personal genome sequencing. Although even Jim McCoy didn't have that on his tricorder.

There is not even the tiniest hint of what is being done with academic collaborators, no posters, no papers, no seminars. Now all we can do is ponder on why the technology has not been launched as planned. Surely it was not pure hype to whip up interest in their last round of funding? Perhaps the technology has suffered issues, there is a lot going into what ONT and others are trying to do; hardware, software, enzymes and just about everything else all need to be designed and tested and any one of these could stop the system from working. Legal wrangling is almost certainly an issue with a huge number of patents in this field and companies like Illumina working on their own Nanopore technology.

So here it is: The blurb reads "Oxford Nanopore chip. Tiny pores 10,000 times smaller than a human hair sit in microscopic holes covering the surface of this speedy chip. DNA is read electronically as it zips through each pore, generating DNA sequence data."

Photo of ONT Nanopore chip in the "Who Am I" gallery at the Science Museum, London.

Sorry the picture and text are so crummy.

I'm sure we'll see the real thing soon, hopefully long before the 23rd century!

PS: The ONT chip is next to an ABI SOLiD instrument. Don't read to much into it being next to a technology that is already obsolete !

Wednesday, 27 February 2013

Fixing Illumina's low diversity problem

Illumina's is the most widely used next-generation sequencing technology, but like all technologies it is not perfect. You'll have to wait for their Nanopore sequencer for perfection! One challenge we have to deal with ever more with Illumina sequencing is the balance and number of libraries in a multiplex-pool. If either of these are wrong, then the final sequencing results can be useless or require significantly more sequencing than would normally be required.

In this post I am going to outline the advice I give to my users.

The low diversity problem: The technology does suffer from a low diversity problem that has been explained over at Pathogenomics. They describe this "fly in the ointment" for Illumina users sequencing amplicons and point to a SEQanswers thread suggesting some work arounds e.g. spike in lots of PhiX. The main problem derives from the need to find distinct clusters on the images, this is done by HiSeq Control Software (HCS) on the fly. The first step is template generation which defines the X,Y coordinates of each cluster on a tile and is the reference for everything else. Template generation currently uses the first four cycles (but can be configured to use more) and this data is analysed after the fourth cycle is complete which users will see as a pause by the instrument. Clusters are found in each of the 16 images (A, C, G & T for four cycles). A Golden Cycle (g) is determined as the one with the most clusters in A&C, this is used as the reference for everything else. By comparing merged A&Cg clusters with A&C and G&T in all cycles a two-colour map of clusters is formed that distinguishes broad clusters from closely neighbouring but distinct ones. However all though there are four bases Illumina instruments only have two lasers, red (A&C) and green (G&T, which I remember by thinking of the colour of a Gordon's gin bottle, green for G&T!) and in each cycle at least one of two nucleotides for each colour channel needs to be read to ensure proper image registration.

See technote_rta_theory_operations.pdf for more detail.

The causes of index read failures: If the index read is unbalanced due to poor mixing or low numbers of indexes then registration can fail due to the low diversity. The balance and number of libraries in a multiplex-pool impacts diversity and low diversity means too many clusters fail dempultiplexing and are thrown away. Illumina provide advice in their sample prep guides (e.g. TruSeq DNA PCR-Free Sample Preparation Guide (15036187 A)) on how best to quantify libraries before pooling and also on how to safely create low-plexity pools (more than 12 is usually safe). I'll now address both of these; Quantifying & Normalising and Deciding Index Plexity in a little more detail.

Quantifying & Normalising: Simply use qPCR quantification to get the very best estimate of nM concentration. Bioanalyser, QuBit and Nanodrop all work but can be very inaccurate when compared to qPCR, it all depends on your libraries. These other methods also quantify molecules that cannot form clusters, such as molecules without both adapters, oligos and nucleotide and can significantly over- or under-estimate pM concentration. A very good QC document is one I found from an Illumina Aisa/Pacific meeting.

Illumina recommends the use of Kapa's KK4824 Library Quantification Kit - Illumina/Universal kit, although now they are selling the Eco systems and their own NuPCR I'd expect an Illumina kit to come along any time soon.

KAPA's is a SYBR-based real-time PCR assay that uses two primers complementary to the ends of TruSeq, Nextera and other library types. During PCR only amplifiable molecules contribute to the CT reading and so the assay is more robust than spectropohotometry or fluorimetry. The method is not without its drawbacks but if care is taken with preparation of the samples and controls then very accurate results are possible.
  • Pipette large volumes - we use a 1:99ul dilution followed by serial 10:90ul dilutions to create a 1:100, 1:1000, 1:10,000 and 1:100,000. 
  • Replicate the dilutions - we make three independent serial dilutions. 
  • Use the controls - we aliquot controls to avoid freeze/thaw. 
  • Include NTC's - vital if you want to see contamination, add them at different stages to see where contamination is coming from.
For library normalisation after sample-prep we have more recently we have moved to a single dilution and only running the 1:10,000 as technical replicates because the results have been so robust. But for clustering we still make the three dilutions above and run the 1:10,000 dilutions as independent replicates. The reason for this is that normalisation can be off a little bit without impacting the experiment but only getting 80M reads when we are expecting 160-180M is embarrassing!
 
All this uses lots of tips and is a pain in the lab but it does give very good results, which are seen in the final barcode balance analysis.

Two words of warning. Firstly total fragment length is required not just the insert size, don't forget the adapters add another 120-130bp. Secondly, if you are using Illumina's TruSeq PCR-free kits then do not use the Bioanalyser to estimate library size as it significantly over-estimates fragment length due to "the presence of certain structural features which would normally be removed if a subsequent PCR-enrichment step were performed". The Illumina sample prep guide has a nice figure illustrating this comparing Bioanalyser and insert size distribution determined by alignment, and suggests using 470bp for 350bp libraries or 670bp for 550bp libraries in the calculations after Kapa qPCR.


Figure 27 from Illumina's TruSeq PCR-free guide

Improving qPCR: A big problem with qPCR as currently implemented is the need to know library size to calculate pM concentration. The assay uses an intercalating dye and is effectively measuring the amount of DNA present rather than the number of molecules. Fragment length is used to calculate number of molecules and if this is wrong then quantitation will also be inaccurate. Ideally we would not bother with the Bioanalyser but we need to understand insert size, this potentially doubles the work we need to do for QT of a library.
 
In my group we discussed a few years ago the use of a TaqMan assay in place of SYBR and other groups have subsequently published methods for this. A TaqMan probe assay counts the number of molecules hydrolysed in each cycle of the qPCR, as such CT gives a direct measurement of the number of molecules present. We're working on an assay to use in the lab as none are available commercially but if someone wants to make one for us please get in touch!
 
Illumina's low-plexity pooling guidelines. Simples!
Deciding Index Plexity or "how many samples should I put in my sequencing pool": Illumina do provide advice in their sample prep guides on how to safely create low-plexity pools but the diagrams can be confusing when you see them for the first time. Usually 12 or more libraries in a pool are "safe" but the mix can go badly wrong. Adding larger numbers of samples to the pool makes this issue disappear, but if a user gets it wrong then a whole lane or flowcell of sequencing can be ruined.
 
Since we started multiplexed sequencing I have been trying to convince users that small pools are the wrong way to go. Most people simply decided to put as many libraries in a single lane as they could and still get the required number of reads at the end of the run, e.g. 4 ChIP-seq samples pooled in 1 SE36bp lane that will generate 160M reads or 40M per sample.
 
I always preferred the mixing of all experimental samples into a super-pool and sequencing as many lanes as required. Users were wary of this because quantification was not robust, but as most have swapped to Kapa qPCR this is no longer a problem. However we still have users giving us pools of three or four samples; which can result in low diversity problems in the index read causing too many perfectly good clusters to be thrown away when they do not demultiplex.
 
Fixing low diversity once and for all: Sequencing is cheap and fast and at least three "simple" fixes appear to be possible. Firstly we can simply use higher numbers of indexes to remove the problem entirely. Secondly we could use a longer cluster definition "read" which would lessen the issue. Thirdly would be to make sure the first four bases of the index read were from random nucleotides. These random bases could be added during adapter synthesis and add four or eight cycles of sequencing to a run.

An added benefit of large "super-pools" being sequenced across multiple lanes and/or flowcells is that failures in the sequencing of any kind can be tolerated as long as most of the data is generated. IF a tumour:normal pair are sequenced as single lanes and one lane fails then no analysis can be performed. If the same pair are indexed and mixed with 7 other pairs and one lane on a flowcell fails copy number analysis would be hardly affected at all. The same principle applies to pretty much any method.
 
Ultimately I'd like to see micro-molar scale synthesis of pairs of i5 and i7 indexes such that almost every sample in a lab is unique. The index pair would be used once an thrown away. For Genomes the cost would be truly negligible, even for RNA-seq with a £60 sample-prep an additional £3-5 on oligos could be tolerated if it removed sample-mixups. 80 pairs of oligos (32 i5 x 48 i7) allow 1536 samples to be run, more than most labs run in a year.

Wednesday, 20 February 2013

Genome partitioning: my moleculo-esque idea

De novo assembly of complex genomes is still harder than many would like.

For Cancer genomes de novo assembly could be the ultimate method for discovery of all somatic events. But the analysis requires long-reads or long-insert reads and this generally eats up lots of DNA in complex methods. The Broad spoke last years AGBT about the "perfect" mix of reads required for using All Paths to assemble a genome (300bp and 3kb I think); and Oxford Nanopore are promising us 100kb reads which would more or less solve the problem, although it may be some time before we can access the technology!

An alternative solution is to sequence a genome the old fashioned way, clone-by-clone, but with NGS. At least one consortium is attempting a BAC tiling-path of the Wheat genome, which is one of the most complex genomes out there. A limitation is the need to sequence some 150-200,000 BAC clones!!!

My "moleulo" idea: At the end of 2011 I was researching alternative digital PCR, and keeping up-to-date with genome-capture methods. While reading around these subjects I had the idea to mix the RainDance emulsion PCR with Illumina's Nextera tagmentation.

Using a set of transposomes that insert unique barcode tags it would be possible to dilute large fragments of DNA such that a single 100kb fragment would be mixed with a single Nextera tag. The resultant transposition library prep would create a set of sequences that all came from the same genomic locus. Et voila a genome ready for two-step assembly; first a local  de novo assembly would stitch together reads from single 100kb fragments, then the long-reads would be stitched together to create the entire genome. Amplifying the DNA first would help and both Moleculo and Population Genetics Technologies (Genome Pooling) have developed their methods to do this.

I had discussed using this on something like the Wheat genome with our Tech Transfer department but they thought it was not practical or protectable. For Wheat I'd ask RainDance to make a library of emulsion droplets from my 200k BAC clones (in the same way they make primer libraries), combine the amplified DNA with the multiplex Nextera droplets and in a single tagmentation get a pretty awesome Wheat genome. It would be possible to use something like Lucigen's NxSeq fosmid prep kit to make a library for the human genome as well.



How does Moleculo work: I still don't know the details, but expect to find out more at AGBT this year where there are several talks on the technology. It appears to be a combination miniaturisation of barcoded-genome library prep in microtitre plate, microfluidics or emulsions and standard NGS. How much DNA it requires as input and how confident you can be about the likelihood of two reads coming from a unique fragment will be the most significant issues for users.

The likelihood of generating reads from a single fragment is going to have something to do with the number of barcodes available and the number of individual reactions performed. The RainDance method I described would allow many 100'000s of tagment reactions to be made so even low-coverage sequence of each one should a robust long-read assembly. There's a whole lot of maths and Poisson distribution statistics that need to be thought!

Illumina tells us more on their website and SEQanswers has a thread on the technology.