Thursday, 21 November 2013

Day 1 at the Royal Society's 2013 Ancient DNA meeting

I spent two very interesting days this week attending the Royal Society’s meeting on Ancient DNA: the first three decades. Recorded audio of the presentations will be available on the Royal Society’s website at some point and the papers will be published in a future issue of Philosophical Transactions B. While at the meeting I made notes during the talks, and I thought that until the recordings have been uploaded to the website these notes might be of interest to those who were unable to attend the meeting. These notes are not intended to provide comprehensive coverage, and I only jotted down items that I personally found of particular interest. My primary focus is on the genealogical applications of DNA testing, and my interests will, therefore not necessarily coincide with those of other researchers. Many of the technical and scientific details of the talks were well outside my expertise. The accuracy of my notes and my interpretation of the lectures is not guaranteed, but I hope that some people might find the information useful.
The Royal Society in Carlton House Terrace, London SW1 - 
the venue for the Ancient DNA meeting.

Full details of the meeting, along with speaker biographies, can be found on the Royal Society’s website. The abstracts for these talks have not been made available on the website though they are all included in the programme which was issued to attendees.

A related satellite meeting is taking place in Buckinghamshire and finishing tomorrow. The speaker’s biographies and the abstracts are available on the website for the this meeting. I was not able to attend this event but I hope that other attendees will provide reports in due course.

Erika Hagelberg, University of Oslo, Norway
Ancient DNA: the first three decades
The first article on ancient DNA was published in 1984. It was a report of the cloning of a small piece of DNA from the skin of an extinct equid (a member of the horse family) that had been preserved in a museum.
The second important ancient DNA paper was on molecular Egyptology.
A lot of the early research centred on Allan Wilson’s lab
In the early days ancient DNA testing was done on the workbench without any protective clothing.
PCR [polymerase chain reaction – a process for amplifying DNA] was introduced in the late 1980s.
The first PCR machine was made with a kettle.
The late 1990s saw the development of standards of authenticity. Hagelberg felt that the new standards stifled research and open discussion.
The big technological advances in recent years have been in bioinformatics, contamination filters and next generation sequencing.
The early studies on ancient DNA (magnolia leaf, an insect embedded in amber) are now not considered very credible. It is also difficult to reproduce these early studies.
The first ancient DNA newsletter was published in 1992.
The limit for ancient DNA was originally thought to be 5000 years.
1 March 1990 Angel of Death newspaper article on the DNA of Mengele. This was the first use of DNA in forensics.
The 1990s also saw the DNA analysis of the remains of the Russian Imperial family. Some people disputed the results.
1994 Dinosaur DNA turned out to be human DNA
1997 Ryk Ward and Chris Stringer publish a paper in Nature in which they outline standards for ancient DNA research
2000 Cooper and Poiner letter in Science. “Do it right or not at all”
Hagelerg said that this was often interpreted as “Do it with me or not at all”.

Christine Keyser, University of Strasbourg, France
Past human populations in Eurasia
Keyser reported on an ancient DNA study of samples obtained from 150 graves in Yakutia  in Northern Siberia.
146 bodies were found. They were frozen at the time of discovery. Genetic data was obtained from 130 bodies.
Optimal ancient DNA is obtained from bone.
Smallpox found in Yakut graves – identified by histology.
Y-chromosome analysis was done using a Y-filer kit (17 Y-STRs). There were 20 different haplotypes. A strong founder effect was found with one haplotype shared by 29 males (46%). They went up to 23 STRs on these samples but found only three differences in the 29 males.
For the mtDNA analysis they tested HVR1 and the coding region. There were 44 different mtDNA haplotypes (n=130) with haplogroups C and D predominating.
IrisPlex and HirisPlex were used to determine hair and eye colour. Six SNPs used to detect eye colour. Brown hair and brown eyes.
SNP testing. N1c1 was the predominant Y-DNA subclade.
Full mtDNA genomes sequenced. D5a2a most common subclade.

Anne Stone, Arizona State University, USA
Impacts of colonisation in the Americas
Anne Stone was invited to speak at the last minute after the scheduled speaker, Ripan Malhi, had to withdraw. Malhi’s talk was to be on the subject of “The evolutionary history of Native Americans”. There is a summary of his planned talk on Science Daily in an article entitled Ancient, modern DNA tell story of first humans in the Americas.

Stone's talk focused on the impacts of colonisation in the Americas.
The initial colonisation of America took place between 18,000 and 25,000 years ago.
The post-Clovis theory of colonisation is dead.
The major part of Stone’s talk focused on the Salesia mission in Tierra del Fuego.
TB was the leading cause of death at the mission. No genetic evidence of TB found in her study.
Targeted enrichment to get full mt genome.
The genetic evidence shows that TB was already in animals in America before humans arrived.
Hershberg et al 2008 paper on the biogeography of M.tuberculosis.
The genetic testing of Native Americans depends on view of individual tribal groups.

Questions from the audience
Q What is the evidence for the pre-Clovis theory?
A The genetic evidence for pre-Clovis colonisation of America is based on signals of expansion. Human coprolite data is also pre-Clovis [coprolite = fossilised poo!].

Helena Malmström, Uppsala University, Sweden
The Neolithic transition in Scandinavia
Farming started 12,000 years ago in the Near East and 7,000 years ago in Northern Europe.
In Scandinavia hunter gatherers and farmers co-existed for a period of about 1000 years.
The hunter gatherers (Pitted Ware complex) and the farmers (Funnel Beaker complex) had different maternal lineages.
Haplogroup U was found at the highest frequency with U4 top of the list.
Autosomal SNP analysis showed that the Neolithic hunter gatherers differ from modern Europeans and were most like Sardinians and Basques.
[DK note: For background see the 2012 Nature News article by Henry Nichols Ancient Swedish farmer came from the Mediterranean and the 2009 paper by Malmström et al.] 

Carles Lalueza-Fox, Institute of Evolutionary Biology (CSIC-UPF), Spain
Neandertal paleogenomics and the El Sidrón cave
This was an excellent and sometimes humorous talk on the exciting findings from El Sidrón cave in Asturias, Spain.
Lalueza-Fox started by sharing a number of illustrations showing how our perception of Neanderthals has changed over time. We now know that they used language, and they lived in family and social groups. The final picture representing the current thinking showed a picture of a Neanderthal mother and child looking not much different from modern humans.
See also the modern reconstruction picture shared by @mjpallen on Twitter.
 Laleuza-Fox took us on a photographic tour of El Sidrón cave. A group of Neanderthal individuals were found in this cave. They had been trapped in the cave after a rock fall and their DNA provides a snapshot in time of a Neanderthal social group.
Complete mtDNA genomes were obtained.  Three different Neanderthal mtDNA haplogroups were found which Laleula-Fox has labelled A B and C. 7/12 were A. 1/12 was B and 4/12 were C. Three adult males had the same mtDNA but the three adult females had different mtDNA. This is indicative of patrilocal reproductive behaviour.
There were cut marks on all the remains – evidence of cannibalism.
Laleuza-Fox et al 2007 paper in Science. Some Neanderthals had red hair

David Reich, Harvard Medical School, USA
Insights into population history from high coverage Neandertal and Denisova genomes
[DK comment: Why do Americans spell Neandertal without an H but pronounce the word as though it does have an H. Why do Brits spell Neanderthal with an H but pronounce it as though it doesn’t have an H?]
This was the highlight of the first day’s talks. It was delivered at breathtaking speed, barely allowing us time to digest the content on the slides. I would have liked to have had a pause button so that I could stop and look at everything again in more detail.
Neanderthal gene flow is about 2%:
1.72% in Europeans
1.89% in East Asians
(Confidence intervals were provided but the slide disappeared to quickly for me to note them.)
Autosomal DNA analysis used a recombination rate of 10cM per 10 generations, 100 cMs per 100 generations. I spotted Graham Coop’s name on this slide but wasn’t sure whether Reich was citing the paper The geography of recent ancestry across Europe 
We now have Neanderthal sequences from three different locations: Croatia, Russia and the Altai Cave in the Altai Mountains in Siberia. This is the cave where Denisovan DNA was found but the latest analyses show that Neanderthals also lived there.
Archaic split 77-114 kya.
There were multiple gene flows.
In the original Denisovan study DNA was extracted from the little finger of a young girl. The samples date back more than 50,000 years. DNA has now also been extracted from a molar.
1.9 fold coverage of genome.
Denisovans are more closedly related to Neanderthals than to humans. Their mtDNA is twice as deep compared to Neanderthals than humans.
Denisovans are closely related to people from New Guinea. New Guineans have 4.6% Denisovan and in addition 2.5% Neanderthal.
2013 paper to be published on Altai Neanderthal found in same cave. Sequencing done at high resolution 52x coverage.
The archaic populations have a very low level of genetic diversity. The Altai Neanderthal are highly inbred.
Reich showed us a number of slides exploring a number of hypotheses he investigated on the relatedness of Denisovans to Neanderthals and humans. He concluded that “Denisovans harbour ancestry from an unknown archaic population unrelated to Neanderthals and modern humans”.
[DK note: This finding was anticipated by Graham Coop in his Haldane’s sieve blog post Thoughts on: The date of interbreeding between Neandertals and modern humans.]
New research has shown that Denisovan DNA is now found in East Asians. See the Cooper and Stringer 2013 paper: Paleontology. Did the Denisovans cross Wallace's Line?
Conclusion: gene flow between diverged humans was common in late Pleistocene and there were five events.

Questions from the audience
Q Does this mean humans copulated with Neanderthals? A Yes!
Q Does this mean humans fancied Neanderthals? A Yes!

Reich’s talk seemed to be the one that was attracting all the interest from the media. Ewen Callaway, the reporter from Nature, was at the conference and he has already written an article for Nature Breaking News which can be found here. There is further coverage from Michael Marshall in New Scientist.

[DK note: The abstract for this paper also mentions Neanderthal X-chromosome ancestry. I don't know if I missed the mention of the X-chromosome in this high-velocity presentation or if it was perhaps not covered. Here is the relevant extract from the abstract: "The average Neandertal ancestry on the X chromosome is about a fifth of that in the rest of the genome. It is known from studies of many species that genetic variations causing hybrid sterility concentrate on chromosome X. This is consistent with Neandertals and modern humans having been on the edge of biological incompatibility when they met and mixed.]

Johannes Krause, University of Tübingen
Ancient pathogen genomics: what we learn from historical diseases
The Black Death killed 30-50% of the population of Europe. It probably originated in China. Yersinia pestis has the biggest diversity in China.
99% of pestis genome sequenced at 30x coverage.
Yersinia pestis MRCA within last 4000 years.
There is nothing in the genome to explain the high mortality rate.

Christina Warinner, University of Oklahoma, USA
A new era in paleomicrobiology: microbiomes
If you go by the number of cells in our body we are 90% bacteria.
The bacteria in our bodies weigh around three pounds.
The bacterial genome is also known as the accessory genome.
There has been a 38-fold increase in the number of known bacteria in the last seven years.
Best estimate before NGS is 500 species of bacteria in mouth. After NGS, 19,000!
You can get lots of DNA from calculus.

[DK note: I'm afraid I was flagging at this point after a 5.15 am start to my day and only four hours' sleep. This talk was highly technical and much of it was over my head. The take-home message from the final talk was that this is an important emerging new field for the study of ancient DNA.]

Update
The recordings of all the lectures from this meeting are now freely available on the Royal Society's website.

See also
My notes from Day 2 at the Royal Society's 2013 Ancient DNA Meeting

© 2013 Debbie Kennett

Sunday, 17 November 2013

Family Tree DNA sale

The Family Tree DNA sale is now on. It is not very easy to find a list of prices on the website so I've copied down all the prices here. The sale ends on 31st December and all tests must be paid for in full by this date. The prices below are shown in US dollars. You can convert the prices into your local currency using one of the many online currency converters. I normally use the XE Currency Converter. Note that the dollar/sterling exchange rate is particularly favourable at present for those of us in the UK!

Basic tests for new customers
Y-DNA 37 markers $119 (usual price $169)
Y-DNA 67 markers $189 (usual price $268)
Y-DNA 111 markers $289 (usual price $359)

mtFull (full mitochondrial sequence) $169 (usual price $199)

Family Finder $99 (US customers also receive a free $100 Restaurant.com gift certificate)

Autosomal DNA Transfer $49 (usual price $69) - this allows people who have tested at 23andMe or AncestryDNA to transfer their autosomal results to FTDNA's Family Finder database

Combination Tests
Family Finder + Y-37 for $218 (usual price $268) 
Family Finder + Y-67 for $288 (usual price $367)
Family Finder + mtFull for $268 (usual price $298)
Y-37 + mtFull for $288 (usual price $366)
Y-67 + mtFull for $358 (usual price $457)
Comprehensive Genome (Family Finder, Y-67 and mtFull) for $457 (usual price $566)

Upgrades
Y-Refine 12 to 37 for $69 (usual price $109)
Y-Refine 12 to 67 for $148 (usual price $319)
Y-Refine 25 to 37 for $35 (usual price $59)
Y-Refine 25 to 67 for $114 (usual price $59)
Y-Refine 37 to 67 for $79 (usual price $109)
Y-Refine 37 to 111 for $188 (usual price $220)
Y-Refine 67 to 111 for $109 (usual price $129)
mtHVR1 to Mega (full mitochondrial sequence) for $149 (usual price $169)

Big Y
This is a new Y-chromosome sequence test for advanced users who are interested in SNP discovery and contributing to our scientific knowledge about the phylogeny of the Y-chromosome. There is an introductory offer on this new test, and It is currently on sale for $495. This test is only available to existing customers. The price will go up to $695 after 1st December. If you have previously taken the Walk Through the Y test you will be eligible for a $50 discount. There should be a voucher that you can use on your personal page. For further information about the Big Y see my earlier blog post on the new Big Y test from Family Tree DNA.

For information on the different types of DNA tests see the beginners' guides in the ISOGG Wiki.

The Y-chromosome sequence interpretation service from YFull.com

This article is for advanced genetic genealogists who have had their Y-chromosome sequenced or who are interested in doing so.

With the forthcoming SNP tsunami, the analysis and interpretation of the Y-chromosome results provided by the various companies will be one of the key determining factors in the success of their products. Fortunately within the genetic genealogy community we have a number of intrepid pioneers who have volunteered to serve as guinea pigs by testing at all the companies so that we will eventually be able to do comparisons between all the products. David Hollister, who runs the Hollister one-name study and is the co-administrator of the Hollister DNA Project, is one of our brave guinea pigs. He has already had his Y-chromosome sequenced with Full Genomes Corporation. He has previously tested with the Genographic Project, and has had STR testing at Family Tree DNA. David is now waiting for his results from the Chromo 2 test from BritainsDNA and the BIG Y test from Family Tree DNA. Another genetic genealogist Itaï Perez has already provided a comprehensive look at the Full Genomes Y-sequencing results in a guest post on CeCe Moore's blog so I see no point in covering the same ground. However, David has recently submitted his Full Genomes data to another service by the name of YFull.com for an alternative interpretation. David was really excited by his results and was so "blown away" by the reports he received from YFull that I asked him if he might be able to share some screenshots so that other genetic genealogists might get a feel for what to expect from this service. David has very kindly agreed and has also obtained the consent of the YFull team for me to publish these screenshots. You will need to click on each image to see larger versions of the screenshots.

This is David's home page on his YFull account. Note that according to YFull there are 41,828 known Y-SNPs and 478 short tandem repeats (Y-STRs).

This report shows David's position on the Y-haplotree and his results for all the SNPs tested on his branch of tree. Separate reports are available for "controversial" SNPs and no calls.

This report provides a list of private and unknown SNPs. 247 private and unknown SNPs were found in David's sequence: 66 were deemed to be of best quality, 10 were of acceptable quality, and 13 were of low quality. For 111 SNPs only one reading could be obtained. A temporary internal ID system is used to identify the private SNPs and they all bear the prefix YFS, an abbreviation for YFull Singleton.

This report shows results for the Indels. Indel is the term used to describe insertions and deletions - positions in the sequence where extra As, Cs, Ts and Gs have been inserted or where they are absent.

There is a handy SNP index that allows you to query your results by SNP name.

Here is the report showing results for the 478 STRs tested.

This pie chart shows the percentage of "good" and "uncertain" alleles. 90.2% of the alleles were classified as "good". Note that next generation sequencing with a read length of 100 bps does not pick up some of the longer STRs in the sequence.

YFull have recently introduced a group feature. There are currently groups available for haplogroups R1a and G2a.

YFull are based in Moscow in Russia. They are currently providing a free service for a limited period, but I understand that they will at some point start charging a small fee. They are able to use data for any Y-chromosome which has been sequenced at a minimum 25X coverage and with a read length of at least 100 base pairs. Data needs to be provided in the form of a BAM file. If you have tested with Full Genomes they will provide you with your BAM file on request. Results are not yet available from Family Tree DNA's BIG Y test but I understand that they will also make the BAM files available. It remains to be seen what level of analysis and interpretation FTDNA will provide.

We can expect the interpretation of Y-chromosome sequencing results to change over time as our knowledge improves, and as more comparative results become available. In the meantime YFull certainly provides an interesting complement to the service provided by Full Genomes. No doubt we can expect other similar services to appear on the scene in the coming months as more sequences become available.

See also
- ISOGG Y-DNA SNP testing chart
- The new Big Y test from Family Tree DNA
- A confusion of SNPs
- A simplified Y-tree and a common standard for Y-DNA haplogroup and SNP nomenclature 

© 2013 Debbie Kennett

Friday, 15 November 2013

A confusion of SNPs

This article is for experienced genetic genealogists and requires a reasonable understanding of SNPs and haplogroups.

The launch of the new Big Y test from Family Tree DNA has brought to light the difficulties in comparing the offerings of the different testing companies. We have a chart in the ISOGG Wiki which compares the various Y-SNP tests on the market but it is clear that we are not always comparing apples with apples. One of the major difficulties relates to the claims by the companies about the number of Y-SNPs on their chip. A SNP is a change or a mutation in the DNA alphabet at a single position on the Y-chromosome (eg, a C changing to a T). There are around 59 million base pairs in the Y-chromosome. However, surprising as it might be in this genomic era, there are still large sections of the Y-chromosome that have not yet been explored. Build 37, the current build of the human genome reference sequence, has only mapped out the positions of around 25 million base pairs  less than half of the Y-chromosome.The discovery of new SNPs is therefore limited to the parts of the Y-chromosome that can be sequenced using current technology. These areas represent just over 40% of the Y-chromosome. In theory, therefore, a SNP could be found on any one of the 25 million bases that can be sequenced.

The exact number of SNPs on the Y-chromosome is not yet known. There is no central resource listing all known SNPs because there is fierce competition and the companies are keen to keep knowledge of the SNPs that they have discovered from their competitors for as long as possible. We therefore have some SNPs that are in the public domain, some unpublished SNPs that are known only to Family Tree DNA/the Genographic Project, some SNPs that are known only to Full Genomes Corporation and some SNPs that are known only to BritainsDNA. To make matters worse all three companies use different naming systems for their SNPs. Full Genomes SNPs are prefixed by the letters FG, and BritainsDNA SNPs bear the prefix S.  I understand from the reports from the Family Tree DNA 2013 Conference that the Genographic Project will be publishing a paper some time in the New Year with the new 2014 Y-SNP tree. It therefore remains to be seen what naming system they will use for their SNPs. There will undoubtedly be considerable overlap in the SNPs offered by the different testing companies but until they release their data or until we have comparative results available we will not be able to work out which SNPs are equivalent (synonymous)  in other words which SNPs occur at the same position but which have been given different names by different companies. For example U106 and S21 are alternative names for a single SNP which defines one of the major branches of the R1b haplogroup.

The problem is well illustrated by the recent developments in R1b-M222, a subclade which predominates in Ireland and Scotland, and is seen in many of the surnames that are associated with the clans reputed to descend from the semi-legendary Irish historical figure Niall of the Nine Hostages.According to the early results from the Chromo 2 testing at BritainsDNA 27 new SNPs have been discovered downstream of M222.3 Yet at the Family Tree DNA Conference last weekend Miguel Vilar from the Genographic Project advised that they have identified 22 SNPS below M222. Do any of the Geno 2.0 SNPs correspond with the SNPs found by BritainsDNA? The answer is we simply do not know. Neither company releases the full raw data that will allow the participant to determine the genome reference position of the SNPs for which he has tested positive so the results from the two companies cannot be compared. Few results are in any case available at present from the Chromo 2 testing. The Genographic Project are presenting the results of their Gathering the Mayo Genes Project at a public event in Castlebar on Sunday so it may be that further information will be forthcoming then.

So where can we find out about SNPs and their position on the Y-DNA haplotree? By far the most important source is the Y-SNP tree maintained by ISOGG - the International Society of Genetic Genealogy. The tree was launched on 10th April 2006. By the end of the year there were 436 SNPs on the tree. By September 2013 there were 3610 SNPs on the ISOGG tree. According to Roberta Estes' report from Day 2 of the FTDNA conference the new 2014 Y-SNP tree, which will be published by the Genographic Project in 2014, will have 6200 SNPS and 1000 branches.This effectively doubles the size of the existing tree and will represent a significant workload for the team of volunteer project administrators who maintain the tree.

However, the ISOGG tree only documents the SNPs whose precise location on the Y-haplotree is known  in other words SNPs that define particular branches of the human family tree on the Y-line. There are thousands more known SNPs. For these SNPs we know that a mutation has been found on the Y-chromosome at the position in question but we do not know if it has any phylogenetic significance, that is, if these SNPs define branches on the Y-tree or if they are unique to the individual.

ISOGG have a SNP index that lists not just the SNPs that are on the haplotree but also those which "are or have been under active investigation and consideration for addition to the Y Haplotree." ISOGG further state that the "SNPs listed here are less than 10% of the currently known SNPs". To supplement the SNP index ISOGG member David Reynolds maintains the ISOGG SNP Compendium Spreadsheet. This was last updated about a month ago and contains a list of 47,680 SNPs which have yet to be added to the ISOGG tree and the SNP index. A small minority of these SNPs are alternative names for previously known SNPs that are already on the tree (for example, some S series SNPs correspond with some of the Z series SNPs that have already been placed on the tree). Most of the rest are SNPs whose position on the Y-chromosome is known but where we do not as yet know where they belong on the Y-tree. David Reynolds reported back in September that he had about another 5000 SNPs to process. He is "curating and combining duplicates" as he goes along so it is a time-consuming process.

There are no doubt many more SNPs that are being published in scientific papers and I don't know if anyone in the genetic genealogy community is currently keeping track of these. In one recent paper uploaded to the ArXiv preprint server two Chinese researchers discovered 25,000 new phylogenetically relevant SNPs.5

Let's now have a look at the offerings of the various testing companies in the light of these numbers. I'm discussing the companies in chronological order based on the dates when their tests were launched. Some companies offer chip-based SNP tests. These tests can only test for previously known SNPs, but the companies can customise the chips to include their own proprietary SNPs for investigation. The new gold standard tests are those which use next-generation sequencing technology. These have the potential to discover thousands of new SNPs.

The Geno 2.0 test from the Genographic Project
The Geno 2.0 test from the Genographic Project was launched in July 2012 and was the first chip test to come on the market with a comprehensive panel of Y-SNPs. The Genographic Consortium published a paper earlier this year with all the technical details of their new GenoChip.6  The supplementary data tell us that the Genographic Project started with "a raw SNP candidate database of approximately 27,500 SNPs" though some of these were duplicates. The original target was to produce a chip with 15,000 SNPs but according to the paper the chip includes around 12,000 SNPs. Customers can download a CSV file with a list of the SNPs. There were 12,059 SNPs in the most recent file that I downloaded for one of my project members. The Genographic Project do not currently provide the genome reference positions of the SNPs on their chip, and it seems likely that this information is being withheld pending publication of the 2014 tree.

The Chromo 2 test from BritainsDNA/ScotlandsDNA
The Chromo 2 test from BritainsDNA/ScotlandsDNA was launched in June 2013. It uses a customised Illumina chip which is advertised as "covering over 15,000 Y chromosome markers, carefully selected to be most informative, and as free from duplication as possible". Only a limited number of results have been released from this test so far, but a flood of results is expected in the next couple of weeks. Customers receive an Excel spreadsheet with a list of all the markers that have been tested. In the one spreadsheet that I've seen there was a list of 14,184 SNPs. Of these, 8,682 SNPs had the S prefix. On the current ISOGG 2013 Y-SNP index the S series SNPs stop at S530. In the list of SNPs that I saw there were 8385 S series SNPs with numbers higher than S530. Many of these SNPs will probably define new branches on the Y-tree but many more could simply be alternative names for currently known SNPs. We do know that the BritainsDNA chip includes SNPs found in the Genomes of the Netherlands Project, and also many SNPs that are likely to be informative for people of British descent. However, BritainsDNA, in common with the Genographic Project, do not publish the genome reference positions of their SNPs. Unless they provide ISOGG with the positions of their SNPs we will have no way of knowing where they fit on the tree and which of their SNPs correspond with those identified by other testing companies.

Full Genomes
Full Genomes is a new start-up company which made a quiet entry onto the market some time towards the end of 2012. They only began advertising their services publicly towards the end of March 2013.7 They currently offer the most comprehensive Y-DNA test on the market covering about 20 to 25 million base pairs representing around 42% of the Y-chromosome. Full Genomes claim to cover 47,000 of the known SNPs on the ISOGG tree and in the ISOGG SNP Compendium. This is after removing "ambiguous results, and synonyms from consideration".8 Around 14 million of the SNPs are reported to be within mappable regions. However, their test is also uncovering many new private SNPs which have not as yet been made public, and the number of new SNPs discovered can be expected to rise as more and more people get tested. At present each testee in one of the common haplogroups can probably expect to find between 25 and 40 private high-quality SNPs. Full Genomes make the raw data available in a BAM file so that customers will have access to the genome reference numbers and can check the ISOGG tree for alternative SNP names as and when new SNPs are placed on the tree.

The Big Y test from Family TreeDNA
The new Big Y test from Family Tree DNA was launched at the weekend at Family Tree DNA's Conference. I've provided preliminary details in a previous blog post. As this is a new test, no results are yet available, and proper comparisons with the other available tests cannot be done. The FTDNA FAQs tell us that the test covers "at least 10 million base-pairs of reliably mapped positions of non-recombining Y-Chromosome", though the exact number of base pairs sequenced has not been disclosed. One conference attendee who spoke to the FTDNA staff was told that "the number of bp [base pairs] analysed will be at least 10 million, but could in some samples go up to 12 million".9 FTDNA claim that their test provides more coverage "than any Y-DNA test on the market".  However, the test is clearly not quite so comprehensive as the Full Genomes test but it does have the virtue of being considerably cheaper which will make testing multiple people within a single subclade a feasible proposition. Confusingly FTDNA claim that the test will cover "nearly 25,000 known SNPs placing you deep on the haplotree". I can only think they've taken their figure of 25,000 known SNPs from the research into the Geno 2.0 chip and that they are seemingly unaware of the ISOGG SNP Compendium Index which, as discussed above, lists over 47,000 SNPs. If they are covering over 10 million SNPs then they will surely test most of the SNPs in the Compendium. Fortunately FTDNA have confirmed that they will make the raw data in the form of BAM files available to their customers so we will eventually be able to make comparisons.

Is next generation sequencing SNP testing for you?
Next generation sequencing is clearly becoming the gold standard for SNP testing. The Genographic Project have announced that they will be introducing a new test within the next seven to 12 months and I would imagine that their new test will use next generation sequencing. No doubt a rival new NGS test is in the works from BritainsDNA too.

The new next generation sequencing Y-SNP tests do have the potential in the long run to be genealogically relevant. There is supposedly a new SNP roughly every one and a half generations. In other words, if there's no SNP found in a son then there will more than likely be a SNP in the grandson. One day the SNPs will effectively allow us to draw complete trees for Y-lines within a genealogical a timeframe. As with any DNA test, a full Y-chromosome SNP test is only useful if you can compare your results with large numbers of other people so that we can work out the chronological order of the more recent SNPs and establish which ones are unique to specific lineages. With the Full Genomes test people in the common haplogroups are reportedly getting between 25 and 40 private SNPs. I imagine the numbers will be pretty similar for the Big Y test from FTDNA.

The numbers of people taking these tests are still relatively small  probably in the hundreds rather than the thousands. Even at $495 a time large-scale testing within a surname project is not going to be a practical proposition. However, if low-hanging SNPs are found that are specific to particular surname lineages then, if these SNPs are added to the a la carte menu, people could test for these single SNPs at $39 a time. STR markers can be used in combination with SNPs to predict who will be positive for which SNP, but ideally you need to be tested to at least 67 markers to make a confident prediction.

The potential problem is that FTDNA are only likely to want to invest money developing single SNPs if there are a reasonable number of people who would be willing to pay for such a test. The more recent the SNPs the fewer people will share them and consequently there will be less chance of the custom SNP tests being developed. FTDNA also only currently have the capacity to offer an additional 2000 custom SNPs. However, they have indicated that they will be re-introducing some form of static deep clade test, probably in the first quarter of 2014, which will be at a much more affordable price. SNPs found in the first phase of the Big Y testing will be candidates for inclusion on these chips so there is possibly some incentive for selected representatives of the various subclades to be tested to ensure that the key new SNPs are included in these tests. Full Genomes have also indicated that they hope to offer single SNPs, and a more economically priced SNP test, but it remains to be seen what they will offer. At the current prices NGS full Y testing is really only for people who wish to contribute to our scientific knowledge and to help delineate all the branches on the Y-tree. No doubt the costs will come down in time. Perhaps in five years or ten years the full Y test will be the norm but we're not there yet.

If you are interested in SNP testing the choice of testing company will be down to the individual and will depend on your budget and your objectives. The ISOGG SNP Testing Chart in the ISOGG Wiki provides a comparison between all the testing companies and is updated as new information becomes available. There will inevitably be new products coming onto the market in the next year with each new test appearing to have a slight advantage over its competitors until the next big thing comes along. I strongly recommend that you join the relevant haplogroup project. The group administrators are all very knowledgeable and will be able to offer good advice. There is a list of Y-DNA haplogroups in the ISOGG Wiki. Most of the projects have associated mailing lists which are currently buzzing with activity, and these will often be the best source of information and commentary.

The SNP tsunami 

The large number of SNPs that will be generated in what has been described as the SNP tsunami will represent a significant challenge for the haplogroup project admins and the citizen scientists who are trying to interpret these data. The new 2014 SNP tree from the Genographic Project, with a mere 6000 or so SNPs, will be something of an irrelevance, and by the time it is published it will be massively out of date, though it will at least lay the foundations for a new nomenclature. The volunteers who maintain the ISOGG tree will have their work cut out to keep up with the new developments. One of the team, David Dowell, has already commented: "It is clear that our processes need to be reorganized and streamlined if we are going to be able to continue to serve the genetic genealogy community and researchers in related disciplines in a timely basis."10

It seems likely that the current confusion will prevail for several months. As one poster on the U106 list has commented, the now infamous quote by Donald Rumsfeld is a very good summary of the current SNP situation:

"There are known knowns; there are things we know that we know.
 There are known unknowns; that is to say, there are things that we now know we  don't know.
 But there are also unknown unknowns – there are things we do not know we don't  know."11
There will be confusion, there will be chaos and there will be competition in the coming months, but from this confusion, chaos and competition many important new discoveries will emerge. I predict that as far as Y-chromosome research is concerned 2014 will be the Year of the SNP.

Updates
Vince Tilroe advises in a comment on Roberta Estes' blog that the 1.5 Y-SNPs per generation was based on the hypothetical presumption that "the entire 60 megabases [60 million bases] of the Y-chromosome could be sequenced. This is not the case by any means, and consequently a more realistic expectation should be closer to 1 Y-SNP per every 4 to 6 generations". Preliminary results from the Full Genomes testing suggest that there is around one Y-SNP every 3 to 4 generations.

Jim Wilson, the Chief Scientist from BritainsDNA, has provided a list of equivalent SNP names for some of the SNPs on the Chromo 2 chip. He has also advised that in due course he will be sharing the genome co-ordinates to allow comparisons with comprehensive Y-chromosome sequences. See CeCe Moore's blog post A list of alternate names for the Y-SNPs from BritainsDNA's Chromo 2 test for further details.

See also
A simplified Y-tree and a common standard for Y-DNA haplogroup and SNP nomenclature
- The Y-chromosome sequence interpretation service from YFull
- YSEQ.net - a new company offering a single SNP testing service

References and notes
1. For further information see the ISOGG Wiki article on the Y-chromosome:  www.isogg.org/wiki/Y_chromosome
2. Moore LT, McEvoy B, Cape E et al. A Y-chromosome signature of hegemony in Gaelic Ireland. American Journal of Human Genetics 2006 78(2): 334–338. Note, however, that this study only used 59 low-resolution STR haplotypes, and many people disagree with the conclusions, both in age and origins.
3. Paterson A. Message posted on the DNA R1b1c7 list. 25 October 2013.
4. Estes R. 2013 Family Tree DNA Conference Day 2DNAeXplained blog, 12 November 2013.
5. Wang C-C, Li H. Discovery of phylogenetic relevant Y-chromosome variants in 1000 Genomes Project data. ArXiv preprint server. Submitted 24 October 2013.
6. Elhaik E, Greenspan E, Staats S et alThe GenoChip: a new tool for genetic anthropologyGenome Biology and Evolution 2013; 5(5): 1021-31.
7. See the thread entitled Full Y chromosome sequencing: Phase III Pilot on the Anthrogenica Forum.
8. Magoon G. Message posted in the R1b-U06 mailing list, 11 November 2013.
9. See the comment thread in the private ISOGG Facebook group at https://www.facebook.com/groups/isogg/permalink/10152015234637922/.
10. Dowell D. ISOGG group gears up for SNP tsunami. Dr D Digs Up Ancestors blog, 13 November 2013.
11. For the background to the quote see the entry for Donald Rumsfeld at Wikiquote: https://en.wikiquote.org/wiki/Donald_Rumsfeld.

© 2013 Debbie Kennett