Monday, 17 September 2012

HuID-seq blog

There has been lots of recent activity around using NGS gene resequencing in the clinic. Although clinical DNA sequencing has been an important tool for several decades the explosion in NGS methods for amplicon resequencing has made it feasible for just about any lab to do. Previous posts on this blog have discussed NGS amplicon methods and some of the tools needed to design amplicons.

It is not easy to say clear which technologie(s) will dominate in the clinical space, nor whether small targeted panels will be preferred over more comprehensive and larger panels, medical exomes or even whole genomes. But it does seem pretty clear that amplicon-sequencing is going to be a very important clinical tool.

Why is patient ID important:
As we are more easily able to sequence not just multiple genes but also multiple patients in s single NGS run it becomes very important to make sure results are not assigned to the wrong patient. Clinical molecular labs spend a huge amount of time and effort on making sure results don’t get mixed up, but I thought the tests themselves could be improved to determine which patient results came from at the same time as the clinical results are being generated. Just add a large enough number of SNP loci to allow patient identification by comparison to a simple blood-based genotyping assay.

SNP-seq for patient ID:
I have been discussing using additional content in amplicon (and other) tests for a year or so, but have never found the time to get in the lab and demonstrate the idea. I asked our Tech-Transfer people about it and they said whilst it was a nice idea there was little that could be protected from an IP perspective. As I am not going to get time to work on it, and as it can’t be protected easily I am hoping this blog will help stimulate discussion and someone will take the idea on board for their research. I call the method HuID-seq.

Comparison to STR profiling: There is already a gold-standard for Human Identification in the STR profiling used in forensic applications. Unfortunately the tests cannot be simply added to an NGS assay. What we need is a level of discrimination so that results cannot be sent to the wrong patient, in theory the HuID-seq could be set at a level significantly lower than forensic STRs. Today 13 STR loci are used in the United States Combined DNA Index System (CODIS) forensic kits. SNPs have lower resolving power and more are likely to be needed but before I get onto that a recent paper deserves a mention. In Biotechniques Bornman et al published an NGS based method that reproduces STR data very nicely (see Short-read, high-throughput sequencing technology for STR genotyping). Potentially this could be added to current tests but I prefer SNPs for a number of reasons.

Why SNPs, which ones and how many: SNPs are a good choice because they are easily assayed by PCR amplification of by in-solution capture methods. SNPs are also already assayed by current NGS methods and perhaps most importantly if coding SNPs are used then RNA-seq data can also be used for HuID-seq. This will be important if array-based gene expression signatures are ported over to NGS. It also means that the HuID-seq method could be used to very good effect in research projects whatever the source of data. It is often important to quality control large experimental datasets to remove duplicate samples or wrongly assigned samples. In the supplementary information for the Nature paper The genomic and transcriptomic architecture of 2,000 breast tumours reveals novel subgroups, the authors presented a novel eQTL-like approach to check that the same patients samples were used for whole genome gene-expression and SNP genotyping arrays. They termed the method BeadArray Diagnostic for Genotype and Expression Relationships (BADGER).

Which ones: There were around 250,000 SNPs present on both Affymetrix and Illumina arrays. 2000 could be used for identification, 100 for ethnicity and 200 for gender. All SNPs would allow mapping of samples onto other SNP or sequence based data, e.g. SNP arrays, ChIP-seq, RNA-seq and exomes or genomes. We looked carefully at which of these SNPs might be used in a HuID-seq method and came up with some requirements.
  • MAF 0.5 (0.4-0.6)
  • Present on most widely used genotyping arrays.
  • Coding SNPs
  • Ability to predict identity
  • Ability to predict ethnicity
  • Ability to predict gender
How many: The number of SNPs is going to be higher than the STR loci, but it need not be very high and the “real-estate” taken up by such a HuID-seq panel need not be large in comparison to an NGS clinical test. SNPs need to be present very broadly in the population to be of any use, ideally using SNPs with MAF0.5 gives any individual a 50:50 chance of being homozygous for one allele or heterozygous. If this 50:50 chance holds true for all SNPs, and if the assay used is perfect then (i.e. no errors in genotyping) then just 4 SNPs give a >90% chance of uniquely identifying an individual. Increasing this to 8 results in a 99.5% and 20 SNPs gives a 99.9999% chance. Choosing a set of SNPS such that there is at least one per chromosome arm leads to a set of about 48. The final number of SNPs used could be determined by the requirement for unique identification, cost or complexity of PCR.

Community cohesion: It makes sense that if this approach is going to be used then everyone should use the same set of SNPs. The best way to do this would be to get people from different clinical and research backgrounds in a room to discuss the why’s and wherefores’ of different SNPs and to recommend a set to the community. If Life Tech, Illumina, Roche, or others get marketing too early on then we are likely to see multiple standards. This is exactly what we have in STR kits with the US and Europe using different sets of loci.

There are already some SNPs being used to QC data by the Broad but the GATK pages have been updated and the older page is now a dead link! LifeTech are launching an 8 SNP Ampliseq based Sample ID kit giving a resolution of 1:5000, you spike the SNP targeting reagents into your assay and go. Illumina are aslo collaborating on forensic products, Dr. Bruce Budowle at the Institute of Applied Genetics, University of North Texas discusses forensic NGS applications on You Tube with the IlluminaInc channel.

It probably has not escaped your notice that a HuID-seq panel could be used for cell line authentication as well. This is something that often gets attention but is easily forgotten by PhD students and post-docs until too late. Making a test part of their ChIP-seq or RNA-seq experiment and comparing back to a reference database would be a simple ad-on.

Thursday, 13 September 2012

My (almost) patent removing the need for users to quantitate before clustering

One of the things I most like about my job is being allowed to think. NGS has provided a very fertile ground for innovation and lots of novel ideas have become products we use regularly today. It is sometimes to easy to forget that any of us can have ideas that are just as good. The missing bit is usually the drive to try and commercialise or patent the idea.

I've tried several times with many ideas but have not been successful yet. In this post I will describe one idea I was sure was going to be the winning one, although in the end someone else had beaten me to it (gues who).

The problem: Illumina NGS systems require very careful quantification of the final sequencing library to generate cluster densities that give the highest yield from a flowcell. Getting QT wrong means over-or under clustering both of which lead to a drop in yield. And over-clustering can mean zero yield or lower quality data. Illumina recommends qPCR (although the method suggested has a few flaws) and many labs have had success wth qPCR, Bioanalyser, QuBit and other QT methods.

Standard Illumina clustering

But wouldn't it be nice if we could just put an unquantified library into a flowcell and always get the correct cluster density? I thought so and came up with a method to do just that.

The idea: basically I wanted to make sure only a certain number of discrete library molecules could hybridise to the flowcell surface for the initial hyb, but allow clustering to proceed normally afterwards. There are currently two oligos in the lawn on the flowcell surface, complementary to the ends of the Illumina adapters. I proposed adding a third “cluster seeding” oligo at a lower concentration. This oligo would have the same sequence as the current flowcell oligos but would be longer and include slightly more of the adapter sequence. Flowcells would be created with a similar lawn of primers such that the seed oligo is present at a spacing consistent with maximal yield and cluster requirements. Additionally the lawn oligos would be blocked by short complementary primers leaving only the 5’ end of the seed oligo unblocked for hybridization with library molecules. A covalently attached library molecule would be produced by polymerase extension. The first round of denaturation would remove the original library molecule and all blocking oligos. Bridge amplification would then proceed as normal.

This method would hopefully allow high, low and/or variable concentration of libraries to be entered into the flowcell without quantification as only a certain fraction would be able to hybridise. The yield of the flowcells would be significantly less variable.

This would remove the need for Illumina customers (you and I) to perform quantification and significantly increase yields from runs.
Seed-oligo clustering idea

Who got there first: Unfortunately when searching through patents to see if anyone else had a similar idea we found this patent US20090226975 by Illumina. They describe a similar method with a clever use of a hairpin oligo so the blocking is removed enzymatically. Fundamentally the same outcome is achievable and quantification can be considered a thing of the past.
US20090226975 figure 3

Why are we all running lots of qPCR: I don't understand why this has not made it out into the current generation of kits. All the users I talk to agree that removing quantification would be a good thing. Even if it is done perfectly it still takes an hour to do so time can be saved.

Come on Illumina, where is it?

Hopefully reading this has spurred you on to move ahead with the good idea you had a couple of months ago and left on the "back burner". I'll follow this post with another one about a completely different idea that is about to be realised by Life Technologies instead!

Ho hum, spending my millions will just have to wait!

TruSight blog

Yesterday Illumina released their clinical research NGS kits called TruSight.
TruSight: There are currently five kits in development, Cancer, Autism, Cardiomyopathy, Inherited disease and Human Gene Mutation Database Exome. Custom kits are sure to follow. Kits are designed to run on MiSeq ad the lab workflow is based on Exome capture protocols recently released as a new product from Illumina.

Combining Nextera with TruSeq (or any other) capture was a smart move (see a previous post on this blog). The simplicity of the library prep much better fits labs than the more complex standard adapter-ligation protocol. In the tests we did in my lab when we first tried the kits we found there was enough library from a 50ng input to sequence a small capture kit (TruSight for instance) and still have enough left over to capture an exome or possibly even sequence the whole genome.

GeneSight: Illumina also announced a partnership with Patners Healthcare on the GeneSight software. Aiming to make this the tool clinical researchers, medical geneticists and pathologists use for analysis. GeneSight is ready for MiSeq data, BaseSpace and the iPad MyGenome app, and is already FDA registered. It should be possible to provide a workflow from sample prep, through sequencing all the way to final analysis and interpretation. The press release mentions that geneSight has been used for over 24000 tests. If this takes off then expect to see graphs similar to ones showing how NGS yield has increased, but for the number of patients reported on by the MiSeq TruSight combo!

Geneinsight has been around since 2005 (this is a good paper describing it) and aims to help with data analysis is several ways. It acts as a repository for data in the form of case histories and variant information. It facilitates clinical reporting. And it will update clinicians as new variants become clinically relevant. This last feature is likely to cause some headaches as patients may need to be told their prognosis has changed based on new information. The GeneInsight network also allows labs to share results and new findings increasing the ease with which variants might be understood to be clinically relevant.

Most of us are going to have to wait though. As is all too often the case with launches like this we have the information on how exciting this is going to be but no release date. A few pilot sites and Illumina’s own CLIA labs will be the first to access the package.

Competition within Illumina: TruSight will go hand-in-glove or head-to-head with the TheraSeq product Illumina recently made a pre-release announcement about. TheraSeq uses the TruSeq custom amplicon approach to target small numbers of clinically relevant genes, see a TheraSeq overview for more details. Currently in development are kits for Non-small cell lung cancer, Advanced metastatic disease and Gastrointestinal stromal tumor. It will work from FFPE tissue and again is likely to run on MiSeq.

If you want to steer Illumina's clinical developments then why not take the survey on the TheraSeq page? One of the interesting questions relates to the size of test panels asking if responders prefer small or large panels, and whether these should be dictated by current standard-of-care or expand beyond it. They also ask for opinions on returning variants of unknown significance. Both hot topics in diagnostic panel discussions.

Also on the TheraSeq page is a link to an overview of Cancer research papers that used Illumina technology.

Summary: Illumina, and all the other NGS companies, know how important clinical is going to be in their future growth. This is unlikely to be just through instrument, consumable and test kit sales. Service provision could be an important revenue generator and this is likely to put Illumina into direct competition with the people they are selling to today. Although according to GenomeWeb Matt Posard said that “Illumina does not intend to offer its TruSight assays as a diagnostic service, because that would compete directly with its customers.

Wednesday, 12 September 2012

Radiogenomics is coming

I posted last week about the emerging field of Immunogenomics. Today I’ve taken a brief look at what is happening in Radiogenomics. Whilst this field is not using NGS in such a comprehensive way I think it can only be a matter of time before it ramps up.

Radiotherapy is an important tool in treating cancer and the impact of genomics on the field was recognised by researchers in Cambridge and Manchester in 2004. Those researchers started the Radiogenomics: Assessment of Polymorphisms for Predicting the Effects of Radiotherapy study (RAPPER) and also helped found the International Radiogenomics Consortium.

The ultimate aim of this consortium is to individualise radiation dose prescription for patients maximising the impact on the tumour whilst minimising normal tissue damage for the patient. They aim to find genomic variants that can help predict how patients will respond to radiotherapy and allow tailoring of treatment. This is somewhat similar to pharmacogenomics approaches used for drugs like Warfarin where SNP genotyping can help establish the correct dose for individual patients. The consortium should make it easier to collect samples for genomic studies and also spur development of methods for radiogenomic research.

The consortium is likely to also learn a lot more about the biology of radiation-induced tissue and DNA damage. Whilst understanding how individuals may respond to radiotherapy is a primary goal, hopefully a better understanding of biology may lead to a list of genes that might be mutated in tumours making them more susceptible to radiotherapy as well.

The Radiogenomics Consotium conducted a GWAS in radiotherapy patients (Independent validation of genes and polymorphisms reported to be associated with radiation toxicity: a prospective analysis study. Lancet Oncol. 2012) to address concerns over how underpowered previous research on late side-effects had been. Late side-effects can have serious impacts on patients and their treatment. This prospective study genotyped 92 SNPs (selected from previous studies) in 1600 breast and prostate cancer patients using the Fluidigm 96.96 Dynamic Arrays. None of the SNPs previously reported to have a significant associations with radiation sensitivity were confirmed. The consortium suggested that the previous associations were “dominated by false-positive associations due to small sample sizes, multiple testing, and the absence of rigorous independent validation attempts in the original studies”.

As the costs of sequencing continue to fall and as associations are found it is likely that NGS will become a more important tool for the consortium. Longitudinal studies of cancer patients can be incredibly revealing and comparison of cancer genome and normal genome with radiotherapy follow up data is likely to yield interesting results.

Thursday, 6 September 2012

Immunogenomics is coming

The immune system is becoming easier to investigate as new methods based on nextgen sequencing are published. I am not an immunologist and the complexities of the immune system for me are stuck back in the days of my undergraduate training. And that was in the 90’s!

Nature and the HudsonAlpha Institute are hosting the first Immunogenomics conference next month bringing together scientists from many disciplines to learn about large-scale immune sequencing projects; epigenetics and the immune system and many other topics. Immunogenomics looks like it is going to make headlines next year.  

There have been several papers describing HLA typing (e.g. Gabreil 2009 & Bentley 2009 using 454 and more recently Wang 2011 using HiSeq) and many groups are working on using next-gen methods to replace older tests.
 
A new product from Sequenta is aiming to make this kind of analysis simple to do for any user. The Lymphosight platform uses a multiplex PCR to amplify the IgH, IgK, TRB, TRG, TRD immune cell receptor loci, allowing each T or B cell to be characterised and counted. Immune cell proliferation in response to disease and other studies might be far easier to carry out using this new kit. With 100’s of millions of reads coming from HiSeq, and eventually Proton, even fairly rare immune cells should be detectable in a high background. 


Lymphosight workflow from Sequenta website

The company discuss a test they ran where sequences associated with a B cell tumour were diluted into a normal background at 1:1,000,000. They got very reproducible, quantitative results and a useful dynamic range that compares well to flow cytometry methods currently being used. They expect Lymphosight to be useful in monitoring of minimal residual disease.

David Haussler Director, Centre for Biomolecular Science and Engineering at UCSC said 'We can read genomes from your immune cells. They adapt throughout your lifetime so they can protect you from diseases. Reading those genomes will be important, and you’re going to hear a lot about them next year.'

I recently visited TRON, a spin out from Mainz University Medical Center, where they are conducting translational research in the field of oncology and immunology. One of their aims is to take personalised immunogenomic markers and turn these into personalised Cancer vaccines. The head of TRON, Ugur Sahin just published a very interesting article in OncoImmunology where they describe using NGS to demonstrate a proof-of-concept for identification of immunogenic tumour mutations that are targetable by individualised vaccines. They analysed a melanoma cell line and found over 500 non-synonymous expressed somatic mutations, one third of which were immunogenic. From these they made long peptides of 27aa length and tested these for immune response. 11 of these immunogenic tumour-specific peptides effectively immunised mice against the Tumorigenic cell line (see figure 1 from their paper below).



I am sure Sequenta are hoping groups like these will be using Lymphosight to do perform their analysis of the Immune repertoire.

Monday, 3 September 2012

How to make Outlook "out-of-office" work for you

Apologies to readers who might have been hoping for some posts over the past few weeks but I have been offline whilst holidaying in France and Spain.

One of my "to-do's" before I left was to respond to the "change your password or you'll be locked out" email from our IT manager. This was one thing that got missed at the end of the frantic Friday afternoon before leaving and subsequently I could not log on, even through webmail.

However, this made for a lovely and uninterrupted holiday and I am sure made it easier to forget about work as there was little point to even try and get online. As a result though I have come back to over 800 emails.

This got me wondering about what I might be writing in my out of office message next time I go away. I think it should go something like this...

"I am currently out-of-the-office on holiday until the 1st of September and will not have any access to email whilst away. If your message really is important then please send it again on the 2nd of September as I will be deleting every email I receive between now and my return to work. Sorry for any inconvenience."

This would certainly make the task I now face much easier!

PS: Normal posting will resume shortly.

Friday, 10 August 2012

Battle of the benchtops part II (you'll need a strong bench for one of these!)

After the furore around the Loman et al paper it is interesting to read another comparison of NGS platforms. Lets face it most of us want to know either what we should be buying next, or if we bought the right thing in the first place.

Comparison papers help.

As do beers at AGBT!

The latest sequencer comparison paper: Mike Quails group at the Sanger published a comparison of PGM, MiSeq and PacBio (interesting choice of the third platform). They sequenced several small genomes that varied massively in GC content. It was interesting to me that these genomes are the routine test genomes for Mikes group, most of us would shudder if a user asked us to sequence something with 20% GC on HiSeq!

Table 1 is excellent reading and should help people in making purchasing decisions. Collecting all this information together needs to be done by each individual institute as prices can vary quite widely. But the table as it stands should allow anyone to make basic comparisons and also see what is missing that they might need to put greater effort into. In the paper they say that although the raw error rate is significantly different for the instruments compared, the affect on SNP calling is negligible given sufficient coverage. 15x appeared fine for the genomes tested. I’d prefer to have seen this in the table as well, to act as a counter to claims around error rates from sales people! They compared most of the things you would want to when deciding what to buy (see the table for everything). The sequencing costs differ significantly per Gb at $500, $1000 and $2000 for MiSeq, PGM 318 and PacBio respectively. This compares to about $50 per GB on HiSeq.


Table 1 from the paper
Most people considering PGM or MiSeq are after a fast sequencer and both will deliver. As we get used to sub-24 hour run times our users will notice how long library prep takes. As costs for sequencing continue to fall we’ll also spend more time questioning the costs of library prep. The paper talked about the push from all companies to make library prep as simple as possible. However there was no mention on the cost of library prep. The genomes sequenced require only ?-? Gb of data and so ??? libraries would be needed per run. At $100 per sample the cost is X?times more than the sequencing. This is still an unmet need of the community, $1 sample prep for $10 genomes.

How did they do the comparison: Genomes sequenced included Bordetella pertussis (68% GC), Salmonella Pullorum (52% GC), Staphylococcus aureus (33% GC) and Plasmodium falciparum (19% GC). They made PCR-free or PCR amplified libraries for MiSeq PE150bp runs, or HiSeq PE75bp lanes allowing a direct comparison of the impact of PCR. Additionally they prepared Nextera libraries from three of the genomes sequenced (Bp, Sa & Pf) and whilst two produced “remarkably even” data the Pf genome was very biased. They made PGM libraries using physical shearing and “Fragmentase” digestion using the Ion Xpress kits and showed both to be comparable. These were run on 316 chips for 65 cycles, generating mean read lengths of 120 base pairs. Standard PacBio libraries were prepared and sequenced using C1 chemistry on multiple SMRT-cells how many?

What did they find: PGM struggled with the very AT rich Pf genome, and the bias appeared to be partly in the library-prep. By tweaking the protocol and swapping the polymerase for a better one they demonstrated a significant improvement in results. Why don’t all companies do this kind of testing before releasing products on us users, using the best polymerase or ligase available can make a huge difference.

Error rates were best for MiSeq, no surprise to Illumina users there. But there was no impact on true-SNP calling with PGM doing best at 15x genome coverage although it did produce more incorrect SNP calls. PGM and MiSeq correctly called 82% and 76% of SNPs and produced 1800 and 1300 incorrect SNP calls respectively. For Illumina MiSeq made more correct SNP calls than HiSeq or GAIIx and Nextera library prep worked as well as the standard protocol. Both MiSeq and PGM’s built-in variant calling was inadequate; MiSeq reporter called 7% and Torrent suite called 1.5% of variants. SNP calling for PacBio was hampered by a lack of tools as most are designed for short-read data.

A word of caution: The paper is out-dated as are all comparisons and the authors are happy to acknowledge this. It takes time to perform an experiment like this, analyse it and finally write it up. C2 chemistry was used for PacBio and a new method has been described for magnetic loading of chips. MiSeq now has 500bp kits available and even more reads. PGM has error rate has improved. MiSeq has an upgrade being rolled out now for more and longer reads. To be fair to the non-Illumina platforms MiSeq is based on a pretty mature technology whilst Ion and PacBio should be given some time to catch-up (and perhaps overtake), some of the issues with the PGM and PacBio might be resolved by evolution.

GenomeWeb had comments from Ion, Illumina and PacBio. Ion and Illumina both said the comparison was fair. Ion clarified this by saying that the data showed what was possible in 2011 but that error rate was now just 0.4%. Whilst IlluminaLoman et al presented.

Mike also spoke to GenomeWeb and said that the same test genomes are still being run and that the results were as valid today as back in 2011. Significant improvements had come from PGM 200 cycle kits and the C2 chemistry for PacBio.

I am confident there will be more of these comparisons in the next few months. Expect at least one AGBT presentation and lots more discussion over beers.

See you at the bar perhaps?

What do celebrities think about science and what do scientists think of celebrity genomes?

Sense About Science is an organisation that tries to provide expert advice on scientific matters to whoever needs it. They monitor the papers and news and produces an annual “Celebrities and Science” round up of the best and worst comments from people in the public eye. The organisation behind Sense About Science has come under some criticism for being pro-GM and a bit radical and certainly not everyone is a fan. But I enjoyed reading through their annual reviews and wanted to share a few of my favourite comments. The best for me was from Nicole ‘Snooki’ Polizzi who said “the oceans were salty because of all the whale sperm”! See the bottom of this post for a selection from the last three years round-ups.

Sense About Science scan many publications looking for comment; of course celebs and politicians don’t always get it wrong but it is far easier to pick up on the crazies out there. There appear to be fewer celebrities who deny evolution or suggest “fossil fuels” aren’t running out, whilst some politicians careers appear to be built on such claims.

If you want to help out then you can sign up or email them with examples of bad science. 

What do Scientists think of celebrity genomes? Jeff Barrett's web page at the Sanger has coverage of a debate on the value of celebrity genomes between Ewan Birney and Paul Flicek. This was part of a series of events at the Sanger institute looking at the relationship between society and personal genomics.

Jeff chaired the debate and before starting the room was evenly split between those who agreed, disagreed or were undecided on the statement “celebrity genomes are a useful contribution to science and society”.

The debate focused around how useful genomes from celebrities were in creating a dialogue between scientists and the public. Paul argued that celebrity genomes are no more important than non-celebrity genomes, so what makes celebrities qualified to speak about genomics? Ewan argued that celebrity genomes have contributed to science, even if only a little. At the end of the debate Jeff asked the audience to judge what impact they thought celebrity genomes had on science, 33% said positive, 62% said negative and 4% were undecided. He also asked the audience if they thought celebrity genomes had had an impact on society, 41% positive, 51% negative and 8% were undecided.

During the debate Paul talked about the impact celebrities can have as patient advocates using Michael J Fox and Parkinsons as an example. Celebrities have as much chance of developing cancer as any of us and as they get their cancer genomes sequenced and see a benefit from the “treatment” they are uniquely placed to talk about the impact in a way that is going to get across to more people than coverage of a Nature paper on the BBC six o’clock news will ever do.

We should be trying to engage with this as much as possible, shouldn’t we?

PS: If you are a celebrity (why wouldn’t they be reading my blog?) and need some advice then help is just a phone call away, call sense about science on +44(0)20 7478 4380. I can’t promise they can say how many reads you’ll need for your next exome sequencing experiment!

PPS: If you want your celebrity genome sequenced there are plenty of labs in LA.


My pick of the best and worst from the annual round-ups.
Positive:

Bonnie Tyler when questioned about trying acupuncture said “I lost some weight but I was also on a more sensible diet at the same time which, if I’m cynical, is more likely the reason for the weight loss.” And Natascha McElhone’s comments about tetanus after a visit to Angola: “It’s completely preventable if you’re inoculated against it.”

Negative:

Heather Mills “meat sits in your colon for 40 years and putrefies, and eventually gives you the illness you die of. And that is a fact.”

Roger Moore “eating foie gras can lead to Alzheimer’s, diabetes and rheumatoid arthritis. In short, eating foie gras is a tasty way of getting terminally ill.” I don’t eat foie gras on compassionate grounds but it is unlikely to be the cause of so many diseases, and I am not sure any of those Sir Patrick listed are actually terminal?

Alex Reid gave out a horrible message about unprotected sex saying “it’s actually very good for a man to have unprotected sex as long as he doesn’t ejaculate” and “semen has a lot of nutrition. A tablespoon of semen has your equivalent of steak eggs, lemons and oranges.” Irresponsible nutter if you ask me!

Julia Sawalha doesn’t get inoculated or take anti-malarials but uses “ homeopathic alternatives, called ‘nosodes’” and said “I’m the only one who never goes down with anything.”

Joanna Lumley, her AbFab co-star put the increase in cancer down to “the growth hormones in the food we eat, that try to make all the chickens, sheep and cows, more productive”.

Sarah Palin who’s autobiography “Going Rogue” says that she “didn’t believe in the theory that human beings — thinking, loving beings — originated from fish that sprouted legs and crawled out of the sea or from monkeys who eventually swung down from the trees.” Yikes how can such strong anti-evolution views be held by someone who (from a UK news coverage perspective) holds some power in the USA?

Michelle Bachman, member of the US House of Representatives and Republican Presidential Candidate, told journalists that a woman had told him her daughter suffered mental retardation after receiving the HPV vaccine, and that this vaccination program has dangerous consequences. What is the likelihood she is a right-wing, pro-christian, pro-guns, anti-abortion Republican?

These last two particularly disturb me. The first highlights how nuts some politicians are. The second because as the UK MMR scare showed, bad science can become mainstream fact and affect us all in a very negative way.

We shouldn’t believe everything we hear in the press, but politicians surely have an obligation to be careful about what they say.

Wednesday, 8 August 2012

What happened to Illumina’s single molecule sequencing or do you remember Solexa’s SMA-seq?

Eight years is a long time in NGS. I recently re-read a 2004 article in Pharmacogenomics 2004, and also found a EBI presentation from Clive Brown and Ewan Birney. Both of these were from a small company based in Cambridgeshire called Solexa. At the time of publication they had only just identified their first alpha-test site and the presentation talked about a prototype instrument ready for the end of 2004. 

Prototpye GA1
Trademarks mentioned in the paper such as SMA-seq and TotalGenotyping have not, I suspect, been heard of by most Illumina sequencing users (including myself). 

The paper describes where Solexa came from (Shankar Balasubramanian and David Klenerman's patents of 1998 spun out of Cambridge University Department of Chemistry). It mentions Solexa's demonstration version of “a system that will allow rapid, base-by-base comparison of genomic DNA sequences” and that this will produce “four or five orders of magnitude improvement over conventional sequencing”. Read lengths of just 25-30bp are proposed, and a nice graph illustrates how just over 80% of the Human genome is uniquely mappable with these incredibly  short reads.

Simon Bennett, business development director of Solexa at the time and author of the Pharmacogenomics paper suggests that Solexa will achieve the $1000 genome within the next ten years. That leaves us two more years to get to $1000 genomes. It does not seem unreasonable that we’ll get there although more discussion today is about the cost of bioinformatics analysis!

What happened to single molecule sequencing: There is an overview of the Solexa Single Molecule ArrayTM technology that the paper suggest can analyse a Human genome in a single experiment. As described there were just 100,000 DNA molecules per cm2 compared to 100M cm2 today. The basic chemistry description is unchanged from current SBS, although only 25 bases were being sequenced at the time of publication.

It is only towards the end of the paper that Solexa’s acquisition of Manteia’s solid surface bridge-amplification technology, this is the clustering we know and love today. Up until this point Illumina had been focusing on single molecule sequencing. Without the acquisition of Manteia perhaps Solexa would have continued to chase single molecule sequencing and ended up like Helicos or Pacific BioSciences. As it stands clustering and SBS chemistry have been the bedrock of next-gen sequencing for the past five years.

Personally I’d bet Illumina are still putting lots of effort into single molecule approaches, and not just by investing in companies like ONT. I’d like to know if it would be possible to sequence single molecules on a HiSeq with a more sensitive camera (massive oversimplification I know)? Imagine 1000M single molecule reads! This might not be what we ultimately use for single-molecule but I think we can be certain there is a lot more coming for next-gen in the next eight years.

PS: Would SOLiD have been the dominant technology if Agencourt had bought Manteia instead? Perhaps we should have a genomics version of Marvel’s “What if” comic books from the 80’s?

PPS: The Illumina history lesson also taught me that we share half our genes with bananas!




Friday, 3 August 2012

Is visual QC of NGS libraries needed anymore?

I have been using the Bioanalyser since its introduction in 1999. Originally intended for QC analysis of total RNA for microarray studies it quickly became a standard tool for many labs. Over the past few years we have run almost as many NGS libraries on DNA 1000 assays as we have RNA chips.

I think we are going to stop using it for all but a small proportion of libraries by next year.

The Bioanalyser has been a great tool for quality control of NGS libraries. Users can clearly see if they have prepared a high-quality library, if there is lots of adapter-dimer present and if the insert size is what they expected. Unforunately running the Bioanalyser is a bit of a pain once you have more than 12 or 24 libraries.

In my lab we are now preparing 24, 48 and 96 libraries in each batch. QC of these has become too much work using current methods so we looked at alternatives. This included the Caliper LabChip GX, Shimazdu MultiNA, Agilent ScreenTape, Qiagen QIAxcel and Advanced Analytical’s Fragment Analyser (see the bottom of this post for a full list of features).





From our analysis of the system features we asked for demonstrations of the Caliper and Advanced Analytical instruments. These two both appeared to give us the throughput and sensitivity we need, both systems worked well and I know of several labs using these instruments very successfully. However we decided not to invest in a high-throughput Bioanalyser.

Why not and what do we want from library QC: most users want sequence results as soon as possible and are happy with some libraries failing so for some the QC is seen as a bar that gets in the way of their science. My lab wants to satisfy all users and return the highest percentage possible of high quality sequencing runs. Generating 40M reads of a poor library is no use to anyone.

With the introduction of 96 and 384 index kits from companies like Bioo Scientific and with Illumina finally catching up with the TruSeq HT kits I think we are ready to ditch gel-based analysis. Instead we will start using a QC pipeline that will use the data from a single lane analysis of up to 96 libraries. We can look at computed insert-size, verify quantification by checking pooling ratios, screen for adapter-dimer or contamination with other genomes and make sure duplication rates are not too high. Even with 96 samples we should get around 1-2M reads each, and some readers of this blog may remember when 1 M reads was considered enough for ChIP-seq analysis, let alone QC! There are also some hints that 1M reads might be acceptable for basic differential gene expression analysis of highly expressed transcripts.

We’ll be slowly retiring the Bioanalyser type analysis of libraries and using the qPCR quantification as a simple QC tool for pass/fail decisions. We might even get to a point that we only quantify the final pool after mixing equal volumes of all 96 libraries, such that cluster density is spot-on. Then we can use the sequence demultiplexing to indicate the actual balance of indexes to re-pool for the final high read number sequencing.

High Throughput Bioanalyser Platform Features
Caliper - Labchip GX
  • High throughput bioanalyser with 96 and 384 well compatibility
  • Asseses RNA quality and gives exact sizing and quantification of DNA fragments.
  • Can analyse 96 samples in less than 1 hour
  • RNA metrics are used to calculate the RGS value (RNA quality score) which has been validated to correlate with the agilent bioanalyser RIN score. This would be beneficial since users are already familiar with a RIN value for assessing RNA quality.
  • Resolution down to 5bp and sensitivity of 0.1 ng/ul
  • Can visualise the results on electropherogram or gel view similar to Agilent 2100.
  • Data can be viewed in tabular form which can be easily exported/uploaded onto our LIMS system.
  • High sensitivity kit also available
  • There is a barcode reader for sample tracking which would be important when running large numbers of samples.
Shimadzu Biotech – MCE 202 MultiNa
  • This is a microchip electrophoresis system for DNA/RNA analysis.
  • Reusable microchips are used which could reduce running and consumable costs.
  • 120 samples can be run simultaneously across 4 separate microchips with 80 seconds per sample processing speed.
  • It can also perform automatic or manual reanalysis of the samples as seen with the agilent bioanalyser and can export the results in a csv. format.
lab901 Agilent - Screentape
  • The Lab901 ScreenTape system is a fully automated system for gel electrophoresis. The ScreenTape instrument loads, separates, images and analyses both DNA and RNA samples. It does this by loading each sample onto a screentape each of which contain 16 microgels which align to built in electrodes and imaging system.
  • Only 1 ul of sample is required and analysis takes 1 minute per sample. It is fully automated with prepacked reagents so there is no gel preparation or chip priming.
  • Different screentapes are available for DNA and RNA analysis.
  • For RNA analysis, quality is displayed as the screentape degradation value (SDV)
Qiagen- QIAxcel system
  • A microcapillary electrophoresis system, which is fully automated and can process up to 96 samples per run. Separation is performed in a capillary of precast gel cartridge which are reusable.
  • Sensitivity of 0.1ng/ul Resolution down to 3-5 bp.
  • Sample consumption is less than 0.1ul, although the minimum sample volume to load for analysis is 10ul.
  • 96 samples can be processed in approximately 1 hour.
  • The data can be viewed as electropherogram or gel images.
Advanced Analytical –Fragment Analyser
  • is a fluorescence-based capillary electrophoresis instrument for both sizing and quantifying nucleic acids (DNA and RNA).
  • Can run either 12 samples or 96 samples at a time
  • The instrument provides space for up to six 96-well plates
  • Can be used to quantify and qualify NGS fragments, RNA, genomic DNA and also for mutation detection, Microsatellite (SSR) analysis.
  • Various capillary lengths can be used, depending on the application, required resolution and desired speed of analysis. Longer arrays provide resolution down to 2 bp for fragments under 300 bp in length. Shorter arrays still provide good resolution with run times as fast as 15 minutes
  • PROSize™ software is used to analyse the data and this can be viewed as a gel view, electropherogram or a results table.
  • The data is exportable and can be linked to the LIMS.