Showing posts with label Other stuff. Show all posts
Showing posts with label Other stuff. Show all posts

Tuesday, 20 September 2016

The future of Illumina according to @chrissyfarr

In yesterdays Fast Company piece Christina Farr (on Twitter) gives a very nice write up of Illumina's history and where they are going with respect to bringing DNA sequencing into the clinic. I really liked the piece and wanted to share my thoughts after reading it with Core-Genomics readers.


Monday, 5 September 2016

Nuclear sharks live for 400 years

A wonderful paper in a recent edition of Science uses radiocarbon dating to show that the Greenland shark can live for up to 400 years - making it the longest lived vertebrate known. See: Eye lens radiocarbon reveals centuries of longevity in the Greenland shark (Somniosus microcephalus).


Friday, 2 September 2016

Sequencing base modifications: going beyond mC and 5hmC

A great new resource was recently brought to my attention on Twitter and there is a paper describing it on the BioRxiv: DNAmod: the DNA modification database. Nearly all of the modified nucleotide sequencing we hear and read about is modifications to Cytosine mostly methyl cytosine and hydroxymethyl cytosine; you may also have heard about 8-oxoG if you are interested in FFPE analysis. All sorts of modified nucleotides occur in nature and may be important in biological processes where they can vary across tissue of an organism, or may just be chemical noise. The modifications are most important when they change the properties of the DNA strand, how is is read, and what might or might not bind to it e.g mC.


Thursday, 1 September 2016

Celebrating 10 years at the CRUK-Cambridge Institute today

Today I have been working for Cancer Research UK for ten years! September 1st 2006 seems like such a short time ago but a huge amount has changed in that time in the world of Genomics. NGS has changed the way we do biology, and is changing the way we do medicine. The original Solexa SBS has been pushed hard by Illumina to give us the $1000 genome, and perhaps just as exciting are the results coming out of Oxford Nanopore's MAP community - this maybe the technology to displace Illumina? What the next ten years will hold is difficult to predict, but today I wanted to focus on the highlights of the last ten years at CRK for me.

CRUK-Cambridge Institue circa early 2006

Thursday, 25 August 2016

Optalysys eco-friendly genomics analysis

The amount of power used in a genome analysis is not something I'd ever thought of until I heard about Optalysys, a company developing optical computing that has the potential to be 90% more energy-efficient and 20X faster than than standard (electronic) compute infrastructure. Read on if you are interested in finding out more, and watch the video below - featuring Prof Heinz Wolff!



Optalysys was originally spun out from the University of Cambridge and the technology needs a lot more explanation that I'll give: briefly they split laser light across liquid crystal grids where each "pixel" can be modulated to encode analogue numerical data in the laser beam, this diffracts forming an interference pattern and a mathematical calculation is performed - all at the speed of light. The beam can be split across many liquid crystals to increase the multiplicity and complexity of mathematical operations performed.

Optalysys and the Earlham Institute in Norwich are collaborating on a project to build hardware/software that will be used for metagenomic analysis. This is a long way from comparing 500 matched tumour and normal genomes in an ICGC project; but if Optalysys can build systems to handle this scale then the huge compute processing tasks might be carried out at a fraction of the current costs and whilst running from a standard mains power supply.

PS: do you remember the Great Egg race as fondly as I do?

Wednesday, 24 August 2016

Upcoming Genomics conferences in the UK

It is almost time for the kick off at Genome Science, probably the best organised academic conference in the UK. It runs from August 30th to September 1st next week and sadly I can't be there (just returned from holidays and too much going on). You can hear from a wide range of speakers in a jam packed agenda. This year it is hosted by the University of Liverpool, and the evening entertainment comes from Beatles Tribute Band “The Cheatles”!

What other conferences are available for Genomics in the UK, and which one should you attend if you too can't make it over to Liverpool? The Wellcome Trust Genome Campus is holding their first Single Cell Genomics conference from September 9th (sold-out I'm afraid). Personally I thought that the London Festival of Genomics was excellent and I've high hopes for the January 2017 meeting. 

Often it is word of mouth that brings a conference to my attention, but there are a couple of resources out there to help.
  • AllSeq maintain a list of conferences.
  • GenomeWeb has a similar list, but it seems less focused than AllSeq.
  • NextGenSeek has a list for 2016, but nothing on the cards for 2017 yet.
  • Nature has an events page (searchable) that lists 50 upcoming NGS conferences.

PS: please do let me know if you've particular recommendations on conferences to attend. And do get in touch with the groups above to list your conference on their sites.

PPS: If you can justify it then the HVP/HUGO Variant Detection Training Course - "Variant Effect Prediction" running from 31st October 2016 is in Heraklion, Crete - a beautiful place to learn!

Thursday, 21 July 2016

Core Genomics is going cor-porate (sort of)

I've just had my five year anniversary of starting the Core Genomics blog! Those five years have whizzed by and NGS technologies have surpassed almost anything I dreamed would have been possible when I started using them in 2007. My blog has also grown beyond anything I dreamed possible and the feedback I've had has been a real motivating factor in keeping up with the writing. It also stimulated my move onto Twitter and I now have multiple accounts: @CIGenomics (me), @CRUKgenomecore (my lab) and @RNA_seq, @Exome_seq (PubMed Twitter bots).

The blog is still running on the Google Blogger site I set up back in 2011 and I feel ready for a change. This will allow me to do a few things I've wanted to do for a while and over the next few months I'll be migrating core-genomics to a new WordPress site: Enseqlopedia.com



Thursday, 14 July 2016

How much time is lost formatting references?

I just completed a grant application and one of the steps required me to list my recent papers in a specific format. This was an electronic submission and I’m sure it could be made much simpler, possibly by working off the DOI or PubMed ID? But this got me thinking about the pain of reformatting references and the reasons we have so many formats in the first place. It took me ten minutes to get references in the required format, and I've spent much longer in the past - all wasted time in a day that is already too full!

Friday, 24 June 2016

I don't want to leave Europe

Brexit sucks...probably. The issue is we don't really know what the vote really means, or even if we'll actually leave the European Union in the next couple of years at all. However one thing cannot be ignored and that is the two-fingered salute to our European colleagues from 52% of the UK voting population that got out of bed on a rainy Thursday.


I am privileged to work in one of the UK's top cancer institutes at the top UK University: the Cancer Research UK Cambridge Institute, a department of the University of Cambridge. The institute, and the University, is an international one with people from all across the globe, many of the staff in my lab have come from outside the UK and they are all great people to work with. I dislike the idea that these people feel insecure about their future because our politicians have done such a crap job on governing the country.

I'd like to keep the international feel so if you're still thinking that working in the UK would be good for you (and your a genomics whiz) then why not check out the job ad for a new Genomics Core Deputy ManagerWe're expanding the lab and putting lots of effort into single-cell genomics service (10X Genomics and Fluidigm C1 right now). I'm looking for a senior scientist of any nationality to help lead the team, with NGS experience, and ideas about single-cell genomics.



You can get more information about the lab on our websiteYou can get more information about the role, and apply on the University of Cambridge website.

Friday, 17 June 2016

Come and work in my lab...

I've just readvertised for someone to my lab as the new Genomics Core Deputy Manager. We're expanding the Genomics core and building single-cell genomics capabilities (we currently have both 10X Genomics and Fluidigm C1). I'm looking for a senior scientist who can help lead the team, who has significant experience of NGS methods and applications; and ideally has an understanding of the challenges single-cell genomics presents. You'll be hands-on helping to define the single-cell genomics services we offer, and build these over the next 12-18 months.



You'll have a real opportunity to make a contribution to the science in our institute and drive single-cell genomics research. The Cancer Research UK Cambridge Institute is a great place to work. It's a department of the University of Cambridge, is one of Europe's top cancer research institutes. We are situated on the Addenbrooke's Biomedical Campus, and are part of both the University of Cambridge School of Clinical Medicine, and the Cambridge Cancer Centre. The Institute focus is high-quality basic and translational cancer research and we have an excellent track record in cancer genomics 123. The majority of data generated by the Genomics Core facility is Next Generation Sequencing, and we support researchers at the Cambridge Institute, as well as nine other University Institutes and Departments within our NGS collaboration.

You can get more information about the lab on our websiteYou can get more information about the role, and apply on the University of Cambridge website.

Thursday, 9 June 2016

Proteomics is starting to rock too...

I don’t usually read Proteomics papers but have been thinking about how we might combine single cell genome and transcriptome sequencing - with Fluidigm’s Helios (CyToF) and have been trying to get more acquainted with Proteomics methods. In doing so I found this excellent paper: Proteomic Biomarker Discovery in 1000 Human Plasma Samples with Mass Spectrometry. The paper is probably a tour-de-force of Proteomics, even if the published results were not stunning, but not being a Proteomecist I’m not sure I’m qualified to say that. It is obvious that the group working on this put a large amount of effort into experimental design ahead of completing the mass-spec work.



Figure from Cominetti et al 2016


Any large project needs to consider design very carefully from considering what factors might need to be controlled for, to deciding what controls to use. The experimental design for the Human proteome paper is illustrated above; they used a controlled-randomised plate layout to remove plate confounding effects for sample origin, gender, age, ethnicity, BMI, blood pressure, glycemic indices, and clinical biochem.

Reducing mass-spec variability with tandem-mass-tags: The key to making the data comparable across what was over 300 mass-spec runs was the use of tandem-mass tags purchased from Thermo Scientific (Rockford, IL, USA), these add specific masses to all proteins in a sample allowing multiplexing of up to 24 samples per run. With a carefully designed experiment it is possible to reduce the impact of run-to-run variability. Much in the same way as we designed projects using multi-sample microarrays, the experimental groups are balanced across mass-spec runs. I’ve learnt a lot more about tandem-mass tagging in Proteomics over the last 18 months after hearing about the tech in an internal seminar. It seems that this approach is going to allow Proteomics researchers to take advantage of the statistical tools developed for gene expression array analysis. The group used a pair of control samples in each run further reducing the impact of technical variability. 304 TMT 6-plex mass-spec runs were performed, with each 6-plex containing two standards, and 4 samples. 1000 patient plasma samples were processed in 19x 96-well plates over a period of just 15 weeks. All sample handling was tracked, although they did not describe their tracking and whether they used a LIMs or not. The paper is a great example of careful experimental design and I thought was one well worth sharing outside the Proteomics community.

Between 150-200 proteins were identified and the authors argue strongly that this was only possible because f the use of TMTs. Label-free mass-spec approaches would have introduced more variability and taken significantly longer (38 weeks by their estimation). However after crunching the numbers only two proteins in Human Plasma had significant correlation with BMI. Both were shown to be associated with obesity.

NGS experimental design: We're lucky to have such a large number of sample barcodes available for NGS experiments. We can usually fit the whole experiment into one library prep plate and a single sequencing pool and remove almost all the confounding technical issues. However this does not mean we should skip careful design of NGS experiments. Taking a little time to discuss the major question(s) being asked, the samples available and the methods we'll use in both wet and dry labs is time very well spent. 

Friday, 24 January 2014

Who won the Core Genomics Christmas Competition?

Congratulations to Tatiana Borodina who has won the Core Genomics Christmas competition and a free MiSeq run courtesy of Illumina. Special mentions to runners up Tim Forshew (Cambridge Institute) and Charles Warden (City of Hope) who were very close to Tatiana. And also to Thomas Rio Frio (Institut Curie) who was the only person to get all 6 baubles right.

Friday, 6 September 2013

Mouse models of Human disease constantly need to be improved

I’ve worked on model-organisms for a long time, originally in Plant research but nowadays I'm more likely to run genomic experiments for groups using Mouse models of cancer. It is impossible to do some experiments in Human patients for a variety of reasons, so we use Mouse models instead. Genetically engineered mice (GEMs) are used in many research programs and offer us the ability to tailor a disease phenotype, as our understanding of the driving events in cancer increases we can build GEMs that carry these same driver mutations; we can even turn the specific mutations on at specific time points to try and recapitulate Human disease.



Tuesday, 23 July 2013

Illumina's next-gen automation solution

Illumina just bought Advanced Liquid Logic. Never heard of them until now, neither had I but I suspect we'll see some pretty cool devices coming soon.

On the ALL website they have the video below, I could not help but watch at 9sec and see PacMan in action, even PacMan conjoining with his twin! ALL makes disposable digital microfluidic devices allowing cost-effective robotic automation; without the robots. Their technology is based around electrowetting and does not use pumps, valves, pipettes or tubes seen on other liquid handling systems.

Monday, 15 July 2013

Managing your researcher profile in the modern age

I have between 16 and 24 publications in various databases and keeping track of these can be more difficult than I think it should be. I've always used PubMed as my primary search tool and have a link to publications by James Hadfield on my blog. However like most of you I have a name that is not so unique, and other James Hadfield's also pop up in my search results (see the end of this post, and feel free to comment if you're another James Hadfield).

I posted previously about the best way to link to a paper and I'm still suggesting the DOI is the thing to use. It can be found by search engines and aggregators making the collection of commentary a little easier. In the same post I also suggested that a unique identifier for an individual would be a big step forward. Well that was also made available recently and in several different forms, so my new quesiton is which ID should you be using?

Who's looking at my papers (or yours): Before I get onto unique researcher IDs I wanted to come back to the issue of how DOI's and other tools allow aggregators to capture content. The newest "killer-app" for me is Altmetric.

They track how papers are viewed and mentioned; in the news, on blogs, on Twitter, etc. The thing you'll probably be adding to you bookmark immediately after reading this post is their free bookmarklet which will give you a report on any paper you happen to be looking at online. Below is an image of the report for one of the recent papers I was involved with. I'd like to see citations tracked and I'm sure we'll see lots more development from the team!

Altmetric

Which profile managing system to use: There are 5 systems you might use and it is hard to say you can easily choose one. However in writing this post I tried to get all 5 up-to-date. Some talk to each other, which helps - but you can't get away from the fact that no-one has time to waste. I'll probably stick with keeping ORCID and GoogleScholar up-to-dte and then link my ORCID ID to my Scopus and ResearcherID accounts.

ORCID: My ORCID ID 0000-0001-9868-4989. The open access one! Started in 2010 ORCID is "an open, non-profit, community-driven effort to create and maintain a registry of unique researcher identifiers and a transparent method of linking research activities and outputs to these identifiers". The registry gives unique IDS to any registered scientists with data being open access. Organisations can also sign up to allow management of staff research outputs. ORCID stands for Open Researcher and Contributor ID.

Google Scholar: Me. Good listing of citations. Good presentation of citation metrics. Want to know how Google Scholar works, then read this.

Scopus:  My Scopus Author ID: 26662876800. Display of citation numbers and link to citations page. Links to web pages mentioning your work and patents referencing it.

ResearcherID: My ResearcherID: A-1874-2013. ResearcherID is a product from Thomson Reuters and provides a solution to the author ambiguity by assigning a unique identifier to researchers. Only articles from Web of Science with citation data are included in the citation calculations. Good presentation of citation metrics.

ResearchGate: Me. Good presentation of citation metrics. Founded by 2 doctors and a computer scientist, Research Gate has been billed as "social networking for scientists". 

Why bother? Given all this information is correctly assigned the the correct authors, at the correct institution and the correct funders then we should be able to get into some very interesting meta-analysis.

Who's collaborative and who's not? Do some institutions or grant funders punch above their weight? If so why, is there something cultural that can be translated to others? Does the journal you publish in impact the coverage your work gets over time? Many other questions can also be asked, although some may not want to see the answers!


My papers in a little bit more detail: Each has a link to its DOI, PbMed and Altmetric stats.

Mohammad Murtaza et al, Non-invasive analysis of acquired resistance to cancer therapy by sequencing of plasma DNA, in Nature 2013 May 2;497(7447):108-12. PMID:23563269 with 8 citations* at the time of writing. Altmetric stats.

Saad Idris et al, The role of high-throughput technologies in clinical cancer genomics, in Expert Rev Mol Diagn. 2013 Mar;13(2):167-81. PMID:23477557 with 2 citations* at the time or writing. Altmetric stats.

Tim Forshew et al, Noninvasive identification and monitoring of cancer mutations by targeted deep sequencing of plasma DNA, in Sci Transl Med. 2012 May 30;4(136):136ra68. PMID:22649089 with 31 citations* at the time or writing. Altmetric stats.

Christina Curtis et al: The genomic and transcriptomic architecture of 2,000 breast tumours reveals novel subgroups, in Nature 486 (7403), 346-352. PMID with 229 citations* at the time or writing. Altmetric stats.

Kelly Holmes et al, Transducin-like enhancer protein 1 mediates estrogen receptor binding and transcriptional activity in breast cancer cells, in Proc Natl Acad Sci U S A. 2012 Feb 21;109(8):2748-53. PMID:21536917 with 15 citations* at the time or writing. Altmetric stats.

Sarah Aldridge & James Hadfield, Introduction to miRNA profiling technologies and cross-platform comparison, in Methods Mol Biol. 2012;822:19-31. PMID:22144189 with 5 citations* at the time or writing. Altmetric stats.

Charlie Massie et al: The androgen receptor fuels prostate cancer by regulating central metabolism and biosynthesis, in EMBO J. 2011 May 20;30(13):2719-33. PMID:21602788 with 56 citations* at the time or writing. Altmetric stats.

Christina Curtis et al: The pitfalls of platform comparison: DNA copy number array technologies assessed, in BMC Genomics. 2009 Dec 8;10:588. PMID:19995423 with 57 citations* at the time or writing. Altmetric stats.

Dominic Schmidt et al: ChIP-seq: using high-throughput sequencing to discover protein-DNA interactions, in Methods. 2009 Jul;48(3):240-8. PMID:19275939 with citations* at the time or writing. Altmetric stats.

Steve Marquadt et al: Additional targets of the Arabidopsis autonomous pathway members, FCA and FY, in J Exp Bot. 2006;57(13):3379-86. PMID:16940039 with X 34 citations* at the time or writing. Altmetric stats.

Raka Mitra et al: A Ca2+/calmodulin-dependent protein kinase required for symbiotic nodule development: Gene identification by transcript-based cloning, in PNAS 2004 Mar 30;101(13):4701-5. PMID:15070781 with 289 citations* at the time or writing. Altmetric stats.

Robert Koebner and James Hadfield: Large-scale mutagenesis directed at specific chromosomes in wheat, in Genome. 2001 Feb;44(1):45-9. PMID:11269355 with 2 citations* at the time or writing. Altmetric stats.

Barbara Jennings et al, A differential PCR assay for the detection of c-erbB 2 amplification used in a prospective study of breast cancer, in Mol Pathol. 1997 Oct;50(5):254-6. PMID:9497915 with 14 citations* at the time of writing. Altmetric stats.


Saturday, 6 July 2013

The GenomeWeb effect

Dan Kobolt wrote a pair of articles about why he suggests you start blogging. In the first he talks about why you should, and perhaps why you would not, start blogging. I'd certainly encourage people to start, it's fun, free and the feedback can be great. Many people leave comments on this blog and I get emails from readers about blogs they'd like to see written. I'm also seeing blogs get +1's now, although I'm not really up to speed with that particular social networking.

The traffic I get to my blog is a real inspiration to keep writing. I regularly meet people who read my blog at conferences and other meetings, although no-one's bought me a beer because they liked it so much! I do keep an eye on my stats and occasionally get a massive spike in readers. Usually this is because GenomeWeb has covered one of my blog posts and the numbers of readers can reach over 1000 a day. I'm sure other bloggers see the same effect on their sites too.

I call this spike "the GenomeWeb effect".

The GenomeWeb effct in action on Core Genomics

Thanks GenomeWeb, I know you can't make them stay but at least you're sending them my way occasionally.

Saturday, 28 April 2012

How do SPRI beads work?

Someone recently asked me, “how do SPRI beads work” and I realized I was not completely sure so I went to find out.

My lab uses kits. Lots and lots of kits; kits for DNA extraction, kits for PCR, kits for NGS library prep and kits for sequencing. Kits rock! But understanding what is going on at each stage of the protocol provided with the kit really helps with troubleshooting and modification. How many people have added 24ml of 100% Ethanol to a bottle of Qiagen’s PE buffer without stopping to ask what is already in the bottle? The contents of the PE bottle are… answer at the bottom of this post.

I hope this post helps you understand SPRI a bit better and think of novel ways to use it. I’d recommend reading A scalable, fully automated process for construction of sequence-ready human exome targeted capture libraries in Genome Biology 2011 for the with-bead methods discussed at the bottom of this post. I am sure we'll all be using their method a lot more in the future!

You are most likely to come across SPRI beads labeled as Ampure XP from Beckman.

First some handling tips:
  • Vortex beads before use.
  • Store beads at the correct temperature.
  • Allow enough time for beads to come to room temperature.
  • Pipetting is critical; use very careful procedures (SOPs) and well calibrated pipettes.
How do SPRI beads work? Solid Phase Reversible Immobilisation beads were developed at the Whitehead Institute (DeAngelis et al 1995) for purification of PCR amplified colonies in the DNA sequencing group. SPRI beads are paramagnetic (magnetic only in a magnetic field) and this prevents them from clumping and falling out of solution. Each bead is made of polystyrene surrounded by a layer of magnetite, which is coated with carboxyl molecules. It is these that reversibly bind DNA in the presence of the “crowding agent” polyethylene glycol (PEG) and salt (20% PEG, 2.5M NaCl is the magic mix). PEG causes the negatively-charged DNA to bind with the carboxyl groups on the bead surface. As the immobilization is dependent on the concentration of PEG and salt in the reaction, the volumetric ratio of beads to DNA is critical.

SPRI is great for low concentration DNA cleanup that is why it is used in so many kits. The reagents are easy to handle and a user can process 96 samples very easily in a standard plate. Alternatively the protocol can be easily automated and tens or hundreds of plates can be run on a robot in a working day. The binding capacity of SPRI beads is huge. 1ul of AmpureXP will bind over 3ug DNA.


This is the typical SPRI protocol from Beckmans website.

Size-selection with SPRI: Again the concentration of PEG in the solution is critical in size-selection protocols so it can help to increase the volume of DNA you are working with by adding 10mM Tris-HCL pH 8 buffer or H2O to make the pipetting easier. The size of fragments eluted from the beads (or that bind in the first place) is determined by the concentration of PEG, and this in turn is determined by the mix of DNA and beads. A 50ul DNA sample plus 50ul of beads will give a SPRI:DNA ratio of 1, as will 5ul pipetting (but much harder to get right). As this ratio is changed the length of fragments binding and/or left in solution also changes, the lower the ratio of SPRI:DNA the higher larger the final fragments will be at elution. Smaller fragments are retained in the buffer usually discarded and so you can get different size ranges from a single sample with multiple purifications. Part of the reason for this effect is that DNA fragment size affects the total charge per molecule with larger DNAs having larger charges; this promotes their electrostatic interaction with the beads and displaces smaller DNA fragments
SPRI size selection from Broad "boot camp"
SPRI size selection from Broad "boot camp"
.
Broad Boot Camp http://www.broadinstitute.org/scientific-community/science/platforms/genome-sequencing/broadillumina-genome-analyzer-boot-camp


With-bead SPRI cleanup: An incredibly neat modification of SPRI clean-up is the “with-bead” method developed by Fisher et al in the 2011 Genome Biology paper. Rather than using SPRI to clean-up discrete steps in a protocol these are integrated into a single reaction tube method, thus reducing the number of liquid transfer steps. After each step DNA is bound to beads by addition of the 20% PEG, 2.5M NaCl buffer, washes are performed as normal with 70% ethanol but the DNA is not eluted and transferred. Rather the master-mix for the next step in the protocol is added directly to the tube. The final DNA product is eluted from beads for further processing e.g. Illumina sequencing. This modification increases DNA yields in Illumina library prep as multiple transfer steps are removed, reducing the amount of DNA lost at each transfer.


How will you use SPRI in your lab?

PS: the Qiagen PE bottle contains 6ml of water.

PPS: If DNA and Carboxyl are both negatively charged what is happening on a SPRI bead? There appears to be a bit of a black hole in the literature I read through about how the SPRI beads actually work at the molecular level. Digging around some more it appears that the PEG can induce a coil-to-globule state change in the DNA with the NaCl helping to reduce the dielectric constant of the solvent, PEG also acts as a charge-shield. These mechanisms may be behind the "crowding" talked about on other sites. Perhaps someone out there can enlighten me?


Refs:
Lis JT. Size fractionation of double-stranded DNA by precipitation with polyethylene glycol Methds Ezymol (1975)
Lis J. Fractionation of DNA fragments by polyethylene glycol induced precipitation. Methds Ezymol (1980)
DeAngelis et al. Solid-phase reversible immobilization for the isolation of PCR products. NAR (1995)
Hawkins et al. DNA purification and isolation using a solid-phase. NAR (1994)

Friday, 27 April 2012

My top ten things to consider when organising a conference

I just organised and hosted an NGS conference for 150 people in Cambridge which was well attended, had interesting talks and received positive feedback. It was a lot of work getting everything organised so I thought I’d put together my thoughts on the whole thing as a reminder of what to do next time and also mention some things I learnt during the process. The tips below should be useful for national meetings of 50-200 people.

If you are going international or much bigger then get more help, and lots more money!

First and foremost in my recommendations for event organisers is plan well and plan early!

Here are my top ten tips, there is some discussion about a few of them further down the page.
  1. Don't forget the science!
  2. What is your meeting about and who is the target audience?
  3. How long will the meeting be and what sort of format will you use?
  4. Who is going to organise it?
  5. Who is going to speak?
  6. Where will the meeting be held and what are the facilities like?
  7. Can you get sponsorship or do you need to charge a registration fee?
  8. How will you advertise and organise registrations?
  9. Who is making name badges?
  10. Get some feedback after the meeting!

1, 2, 3: The science and the audience are your most important considerations. Make sure you have a good understanding of what the audience are expecting to get from the meeting. Your own experience of good (and bad) conferences will help. Understanding the focus of the meeting will also help determine the probable length of your meeting. A lot can be accomplished in half a day, more than one and a half might mean two nights accommodation and that can put people off coming.

4, 5, 6: Organisation is key, get this wrong and the meeting might not be as successful as you'd hoped. Are you going to organise it yourself or will someone in your lab, or you might even have an administrator who can help? Will the meeting be in your host institution or somewhere else, what sort of rooms do you have available and is there a pub nearby for afterwards!

7, 8, 9: Meetings cost money and people need to be told about them. If you can get sponsorship then do, many companies would love the opportunity for a captive audience. Charge more for a stand or the chance to give a talk, between £100-£1000 for a stand and as high as you can for a talk. You could also offer companies the chance to sponsor an academic speaker and cover their flights and accommodation, this goes down better with audiences. Registration fees can cost more to administer than they raise in revenue, speak to your finance officer about them first. We used Survey Monkey for registrations, users can answer questions about session preferences and food requirements and everything is delivered back in a simple spreadsheet format with email addresses for a confirmation from you to your attendees. Don't underestimate name badges. Making 200 yourself without knowing what a "mail merge" is would be hard work.

10: Get some feedback. I used Survey Monkey again to ask several questions about the 3rd CRI/Sanger NGS workshop (all the results are here). First and foremost was did people think the meeting was good (91% said yes)? Would they come again (90% said yes)? I also asked about their background and experience, as well as the aspects of the meeting they liked best to help shape the next years meeting. Leaving space for some free text comments was good as well, some people gave very useful constructive criticism. 

Other things: I used a spreadsheet in Excel to keep track of the meeting, what the times were for each session, who the speaker was, had they confirmed, did I have a title, what were their contact details, etc. You can get a copy here.

A hard thing to get right is how much to over-book your meeting or under-cater it. You can bet there will be no-shows on the day and a registration fee will help prevent speculative registrations. But inevitably some people won't be able to make it. I would aim for filing the room allocated next time but only catering for 75% of the numbers to keep costs reasonable.


PS: Remember to ask people to switch of their mobiles. I was at a meeting where a phone rang loudly for a while before the current speaker realised it was theirs!

Thursday, 26 April 2012

3rd CRI/Sanger NGS workshop

Over the past two days I have hosted the 3rd CRI/Sanger NGS workshop at CRI. I hosted the first of these in 2009 and the demand for information about NGS technologies and applications appears to be greater today than it was back then. We had just under 150 people attend, and had sessions on RNA-seq, Exomes and Amplicons as well as breakouts covering library prep, bioinformatics, exomes and amplicons. Our keynote speaker was Jan Korbel from EMBL.

I wanted to record my notes from the meeting as a summary of what we did for people to refer back to. I'd also hope to encourage you to arrange similar meetings in your institutes. Bringing people together to talk about technologies is an incredibly powerful way to get new collaborations going and learn about the latest methods.
Let me know if you organise a meeting yourself.

I will follow up this post with some things I learnt about organising a scientific meeting. Hopefully this will help you with yours.

3rd CRI/Sanger NGS workshop notes:

RNA-seq session:
Mike Gilchrist: from NIMR presented on his groups work in Xenopus development. He strongly suggested that we should all be doing careful experimental design and consider that variation within biological groups needs to be understood as methodological and analytical biases are still present in modern methods. mRNA-seq is biased due to transcript length, library prep and sample collection, which needs to be very carefully controlled (SOPs and record everything).
Karim Sorefan: from UEA presented a fantastic demonstration of the biases in smallRNA sequencing. They used a degenerate 9nt oligo with 250,000 possible sequences as the input to a library prep and saw heavy bias (up to 20,000 reads from some sequences, where only 15 expected, similar issues with a larger 21nt 1014 sequences) mainly coming from RNA-ligation bias. smallRNAs are heavily biased due to RNA-ligase sequence preferences, some RNA-ligase bias comes from the impact of 20 structure on ligation (there are other factors as well). He presented a modified adapter approach where the 4bp at the 3’ or 5’ ends of the adapters are degenerate, and very significantly improved bias to almost the theoretical optimum.
Neil Ward: from Illumina talked about advances in Illumina library preparation and the Nextera workflow, which is rapid (90 mins) and requires vanishingly small quantities of DNA (50ng) without any shearing and uses a dual-indexing approach to allow 96 indexes from just 20 PCR primers (8F and 12R, possibly scalable to 384 with just 40 primers). Illumina provide a tool called “Experiment Manager” for handling of sample data to make sure the sequencing run can be demultiplexed bioinformatically. Nextera does have some bias (as with all methods) but is very comparable to other methods. He showed some amplicon resequencing application (drop off of the ends of PCR products at about 50bp), with good coverage uniformity and demonstrated the ability to detect causative mutations.

Exome session:
Patrick Tarpey: Working in the Cancer Genome Project team at Sanger was talking about his work on Exomes sequencing. Refreshed everyone on how mutations can be acquired somatically over time, that there are differences between Drivers and Passengers and that exome sequencing can help to detect  drivers where mutations are recurrent e.g. SFB31. They are assaying substitutions, amplifications and small InDels (rearrangements and epigenetic not currently assayable). He presented a BrCa ER+/- T:N paired exome analysis using Agilent SureSelect on HiSeq with Caveman (substitutions) and Pindel (InDel) analysis. About 1/3rd of data is PCR duplicates or off-target so they are currently aiming for 10Gb with 5-6Gb on-target for analysis, ~80% bp at 30x coverage. Validation is an issue does it need to be orthogonal or not? So far they have identified about 7000 somatic mutation in the BrCa screen (averaging around 10 indels and 50 substitutions) however some mutator phenotype samples have very large numbers of substitutions (50-200). An interesting observation was that TpC (Thymine precedes Cytosine) dinucleotides are prone to C>T/C/A substitutions (why is currently unknown). They found 9 novel cancer genes including two oncogenes (AKT2 & TBX3). And noted that many of the mutated genes in the BrCa screen abrogate JUN kinase signalling (defective in about 50% of BrCa in the study).
David Buck: Oxford Genomics Centre, spoke about their comparison of exome technologies (Agilent vs Illumina). His is a relatively large lab (5 HiSeqs) with automated library prep on Biomek FX. He talked about how exomes can be useful but have limitations (SV-seq is not possible, not clear about impact on complex disease analysis). They did not see a lot of difference in the exome products and he talked about the arms race between Agilent and Illumina. They recently encountered issues in a 650 exome project and saw drop-out of GC rich regions, almost certainly due to a wrong pH in NaOH from acidification during shipment. They had seen issues with SureSelect as well. All kits can go wrong make sure you run careful QC of the data before passing it through to a secondary analysis pipeline.
Paul Coupland (AGBT poster link): Talked about advances in genome sequencing. He discussed nanopores and some of the issues with them, focusing on translocation speed. And there was some discussion about whether we would realy get to 50x genome in 15 minsSMRT-cell sequencing. Briefly they are blunt-end ligating DNA to SMRT-bell adapters to create circular molecules for sequencing, read lengths are up to 5kb but made of sub-reads (single pass reads) averaging 1-2kb, 15% error rate on single pass sequencing but <1% on consensus called sequence, vast majority of errors are insertions, variability is often down to library prep rather than sequencing, improved library prep methods are needed. He completed a P falciparum genome project in 5 days DNA on 21 SMRT-cells for 80x genome coverage, long reads are really helpful in genome assembly. Paul also also mentioned the Ion Torrent technology as well as 454, Visigen, Bionanomatrix, Visigen, GnuBio, GeniaChip, etc.

Amplicons and clinical sequencing:
Tim Forshew: Tim is a post-doc in Nitzan Rosenfeld's group here at CRI. He is working on analysis of circulating tumour DNA (ctDNA) and its utility for tumour monitoring in patients via blood. ctDNA is dilute, heavily fragmented and the fraction of ctDNA in plasma is very low. The group are developing a method called "TAM-seq" which uses locus-specific primers tailed with short tag sequences, followed by a secondary PCR  from these to add NGS adapters and patient specific barcodes. They are pre-amplifying a pool of targets, then using Fluidigm to allow single-plex PCR for specificity and generating 2304 PCRs on one AccessArray. PCR products are recovered for secondary PCR and the sequenced. Till now they have sequenced 47 OvCa (TP53, PTEN, EGFR*, PIK3CA*, KRAS*, BRAF* *hotspots only), with amplicons of 150-200bp (good for FFPE), from 30M reads on GAIIx for 96 samples they achieve >3000x coverage. Tim presented an experiment testing the limits of sensitivity using a series of known SNPs in five individuals mixed to produce a sample with 2-500 copies per SNP and they looked at allele frequency for quantitation. Sanger-sequencing confirmation of specific TP53 mutations was followed-up by digital PCR to showed these were detectable at low frequency in the pool. Fluidigm NGS found all TP53 mutations and additionally detected an unknown EGFR mutation present at 6% frequency in the patient, this mutation was not present in the initial biopsy of the right ovary. Tim performed some "digital Sanger sequencing" to confirm EGFR mutations, and looking in more detail the biopsies of Omentum showed the mutation whilst it was missing in both ovaries. The cost of the assay is around $25 per sample for duplicate library prep and sequencing of clinical samples, they can find mutations as low as 2% (probably lower as sequencing depth and technology gets better), and are now looking at the dynamics of circulating tumour DNA. TAM-seq and digital PCR are useful monitoring tools and can detect response or relapse weeks or even months earlier than current methods, it is also possible to track multiple mutations in a patient as complex-biomarker assays probably allowing better analysis of tumour heterogeneity and evolution during treatment. He summarised this as a personalised liquid biopsy!

Christohper Watson: Chris works at St James's clinical molecular genetics lab which runs a series of NGS clinical tests biweekly to meet turn-around-time (TAT), so far they have issued over 1500 NGS based reports. There's was the first lab in the UK to move to NGS testing, see Morgan et al paper in Human Mutation. They perform triplicate lrPCR amplification for each test and pool before shearing (Covaris) for a single library prep (Beckman SPRIworks) for Illumina NGS, aiming for reports to be automated as much as possible, still doing Sanger sequencing where not enough individuals to warrant NGS approaches, lrPCRs are pooled for library prep by shearing (Covaris) and standard Illumina library prep on Beckman SPRIworks robot (moving to Nextera), they run SE100bp sequencing reads (50x coverage) and are using NextGENe from SoftGenetics for data analysis that gives mutation calls in standard nomenclature. Currently all variants are confirmed by Sanger-seq (100% concordance) but this is under review as the quality of NGS may make this unnecessary and other NGS platforms may be used instead. The lab have seen a 40% reduction in costs and 50% reduction in hands on time their process is CPA accredited. They see clinicians moving from sequential analysis of genes to testing all at once improving patient care and now 50% of the workload at Leeds is NGS. Test panels cost £530 each per patient. Chirs discussed some challenges (staff retraining, lrPCR design, bioinformatics, maximising run capacity), and they are probably moving to exomes or mid-range panels, the "medical exome". On thing Chris mentioned is that in the UK TAT for BRCA is required to be under 40days, which was a lot longer than I thought.

Howard Martin: Is a clinical scientist working at EASIH at Addenbrokes Regional genetics lab where they are still doing most clinical work with Sanger-seq offered on a per gene basis which can take many months to get to a clinically actionable outcome. They initially used 454 but this has fallen very much out of favour, Howard has led the testing and introduction of the Ion Torrent platform, he had the first PGM in the UK and likes it for a number of reasons; very flexible, fast run time, cheap to run, scalable run formats, etc. They are getting 60Mb from 314 chips. They have seen all the improvements promised by Ion Torrent and are waiting for 400bp reads. He presented an HLA typing project for bone marrow transplantation where they need to decode 1000 genotypes for the DRB1 locus. He is also doing HIV sequencing to identify sub-populations in individuals from a 4kb lrPCR. They are also concatenating Sanger PCRs as the input to fragment library preps to make use of the up-stream sample prep workflows for Sanger-seq, costs are good data quality is good with a very fast TAT.
 
Graeme Black: Graeme is a clinician at the St Mary's hospital in Manchester. He is trying to work out the best way to solve the clinical problem rather than looking specifically at NGS technologies. He needs to solve a problem now in deciphering genetic heterogeneity. The NGS is done in clinically accredited labs. He is working primarily on retinal dystrophies that are highly heterogeneous conditions involvinglem now in deciphering genetic heterogeneity. The NGS is done in clinically accredited labs. He is working primarily on retinal dystrophies that are highly heterogeneous conditions involving around 200 genes. With Sanger-based methods clinicians need to decide which tests to run as many patients are phenotypically identical, they cherry-pick a few genes for Sanger sequencing but only get about 45% success rates. The FDA looking to accredit gene therapies for retinal disorders already. Graeme would like success rates on tests to be much higher, complex testing should be possible and universally applied to removes inequality of access (~20% patients get tested). They are now offering a 105 gene test (Agilent SureSelect on 5500 SOLiD) and Graeme presented data from 50 patients. A big issue is that the NHS finds it really difficult to keep up with the rapid change in technology and information, the NGS lab has doubled the computing capacity of the whole NHS trust!!!. They have pipelined the analysis to allow variant calling and validation for report writing and are finding 5-10 variants per patient, results from the 50 patients presented showed 22 highly pathogenic mutations and they were able to report back valuable information to patients. He is thinking very hard about the reports that go back to clinicians as carrier status is likely to be important. 8/16 previously negative patients got actionable results and he saw a 60-65% success rate from the NGS test. Current single-gene Sanger £400, complex 105 gene NGS cost £900, NHS spends more but patients get better results www.mangen.co.uk.


Bioinformatics for NGS:
Guy Cochrane, Gord Brown, Simon Andrews, John Marioni: This was one of the breakout sessions where we had very short presentations from a panel followed by 40 minutes discussion. Guy: Is a group leader at the EBI who are the major provider of Bioinformatics services in Europe; databases, web tools, genomes, expression, proteins etc, etc, etc. He spoke about the CRAM project to compress genome data currently we use 15-20 bits per base, CRAM in lossless is 3-4 b/b and it should be possible to get much lower with tools like reference based compression. Gord: Is a staff bioinformatician in Jason Carroll's research group at CRI. He spoke about replication in NGS experiments and its absolute necessity. Why have we seemingly forgotten the lessons learnt with micorarrrays? Careful experimental design and replication are required to get statistically meaningful data from experiments. Biological variability needs to be understood even if technical variability is low. Gord showed a knock-down experiment with singletons (looked OK) but replication showed how unreliable the initial data was. Reasons why people don’t replicate; cost, “if you can’t afford to do good science is it OK to do cheap bad science?”, sample collection issues, time, etc. John: Is a group leader at the EBI who is working on RNA-seq. He talked about the fact that it is almost impossible to give prescribed guidelines on analysis, the approach depends on your question, samples, experiment, etc. Things are changing a lot. Sequencing one sample at high depth is probably not as good as sequencing more replicates at lower depth (with similar experimental costs). There are many read mapping tools (genome, transcriptome fast but you could miss stuff, tools, de novo gene models 60 tools at last count) but it is difficult to choose a “best” method. Many are using RPKM for normalisation but it is still not clear how best to normalise and quantify. There is a lack of an up-to-date gold standard data set (look out for SEQC). john thought that differential expression detection does seem to be a reasonably solved problem, but more effort is needed for calling SNPs in the context of allele-specific expression. He pointed out that there is no VCF-like format for RNA-seq data that might be useful to store variation data in RNA-seq experiments. Simon: Is a bioinformatician at the Babraham Institute and wrote the FastQC package. It should be clear that we need to QC data before we put effort into downstream analysis! Sequencers produce some run QC but this may not be very useful for your samples. Library QC is also still important. Why bother? You can verify your data is high-quality, not contaminated and not (overly) biased. This analysis is also an opportunity to think about what you can do to improve any issues before starting the real analysis. You can also discard data, painful though that may be! Your provider should give you a QC report. Look at Q-scores, sequence composition (GC, RNA-seq has a bias in the first 7-10 nucleotides), trim off adapters, check species, duplication rates, aberrant reads, (FastQC, PrinSeq,QRQC, FastX, RNA-seq QC FastQScreen, TrimGalore).

An issue that was discussed in this session was "barcode-bleeding" where barcodes appear to switch between samples, not sure what is going on, are we confident that barcoding has well understood biases?

Our Keynote Presentation:
Jan Korbel: Jan is a groupl leader at the EMBL lab in Heidelberg he was previously in Mike Snyders group in Yale where he published one of the landmark papers in Human genomic structural variation. Jans talked about phenotypic impact of SV in human disease, rare pathogenic SVs in very small numbers of people e.g. Downs muscular dystrophy, common SVs in common traits psoriasis cancer and cancer-specific somatic SVs e.g. PrCa TMPRSS2/ERG, leukaemia BCR/ABL.

His groups interest is deciphering what SVs are doing in Human disease. Using SV mapping from NGS data Korebel et al 2007 Science They have developed computational approaches to improve the quality of SV detection, read-depth, split-reads, assembly. He is now leading the 1000 genomes SV group. They produced the first population level map of SVs and looked at the functional impact in GX data. A strong interest is in trying to elucidate the biological mechanisms behind SV formation in the Human genome, non-allelic homologous recombination (NAHR) 23%, mobile element insertion (MEI) 24%, variable number of simple tandem repeats (VNTR) 5%, non homologous rearrangement (NH, NHEJ) 48% (%ages from 1000 genomes data). Some regions of the genome appear to be SV hotpsots and are often close to telomeric and centromeric ends.

Cancer is a disease of the genome: it is often very easy to see lots of SV in karyotopes, every cancer is different, but what is causal or consequential? The Korbel group are working on the ICGC PedBrain project; in medulloblastoma there are only about 1-17 SNVs per patient with very poor prognosis. He presented their work on the discovery of “circular” (to be proven) double minute chromosomes. In teir first patient thay have confirmed all inter-chromosomal connections by PCR as present in the Tumour only. In a 2nd patient chr15 was massively amplified and rearranged at same time Stephens et al at the Sanger Institute described “chromothripsis” in 2-3% of Cancers. Jan asked the question “is the TP53 mutation in li fraumeni syndrome driving Chromothripsis?” They used Affy SNP6 and TP53 sequencing to look and saw a clear link between SHH and TP53 mutation in Chromothripsis in medullobastoma. The data is suggestive that chromothripsis is an early, possibly initiating event in some medulloblastomas and these cancers do not follow the textbook progressive accumulation of mutations model of cancer development and that it is primarily driven by NHEJ.

TP53 status is actionable and surveillance of patients with MRI, mammography increases survival and personalised treatments may be required as exposure to radiotherapy, etc could be devastating to these patients.