Showing posts with label Genographic Project. Show all posts
Showing posts with label Genographic Project. Show all posts

Sunday, 2 June 2019

The end of public participation in the Genographic Project

It is the end of an era. The National Geographic Genographic Project has announced that the public participation phase of the project has been closed as of 31st May 2019.  It is no longer possible to order a Genographic kit, but existing orders will be fulfilled within a limit timeframe with the date varying depending on which kit was ordered. There is further information on the Genographic Project website:

https://web.archive.org/web/20190627191303/https://genographic.nationalgeographic.com/ (retrieved from the Wayback Machine)

The Genographic Project has provided a detailed set of FAQs:

https://web.archive.org/web/20190716185315/https://genographic.nationalgeographic.com/faq/sales-shutdown-previous-kits/ (retrieved from the Wayback Machine)

As of today's date, the Genographic Project has sold 997,222 kits in 140 countries.

There are no doubt many kits still waiting to be returned and it's possible that the project will eventually pass the one million milestone.

This was an almost inevitable development after Rupert Murdoch bought out the media arm of the National Geographic and ended its not-for-profit status. The new for-profit arm was re-named as National Geographic Partners and was went into partnership with Disney in March this year. The National Geographic Society continues to operate as a non-profit organisation.

The Genographic Project was not without controversy. See for example the essay The brave new era of human genetics by Hans-Jurgen Bandelt, Yong-Gang Yao, Martin Richards and Antonio Salas published in 2008. The Native American researcher Kim Tallbear published a critique Narratives of race and indigeneity in the Genographic Project in 2007. Many population geneticists were critical of the fancy Y-DNA and mtDNA haplogroup stories provided as customer reports. Ancient DNA testing has now shown that we cannot use the DNA of living people to make inferences about past populations.

However, many genealogists first discovered the joys of genetic genealogy by testing at the Genographic Project. After transferring their DNA results to FamilyTreeDNA many people were then inspired to start their own surname projects, haplogroup projects and geographical projects.

The Genographic Project collected DNA from nearly 100,000 people from indigenous populations around the world. I understand they were waiting for the costs of whole genome sequencing to come down before starting to analyse all the data. This is a valuable resource and the scientific research will continue so we can look forward to many more interesting publications.

Anyone who has tested at the Genographic Project can transfer their data to the FamilyTreeDNA database:


Note, however, that Helix kits, which were sold exclusively in the US, cannot be transferred.

Genographic transfers will have the kit number prefixed by the letter N. Judging by the kit numbers in my projects at FTDNA, well over 200,000 people have already transferred their Genographic results to FTDNA.

When transferring to FamilyTreeDNA you need to be aware that if you participate in relative matching the company is now automatically opting all customers into Law Enforcement Matching. This means that DNA profiles uploaded by law enforcement agencies in the US and their representatives can access your name, your e-mail address and the amount of DNA you share with the the law enforcement kits. Law enforcement matching is not restricted to US citizens but applies to the entire database regardless of country of residence. If you wish to opt out of Law Enforcement Matching you can do so from the Privacy and Sharing Page. If you wish to understand more about these issues you can read my article for Forensic Science International on Using genetic genealogy databases in missing persons cases and to develop suspect leads  in violent crimes.

With thanks to Mats Ahlgren and Paul R Smith in the ISOGG Facebook group. See also Paul's blog post National Geographic Geno Project DNA ending.

Further reading
Genographic Project prepares to shut down consumer database by Roberta Estes, DNAeXplained

Wednesday, 23 November 2016

Exome testing combined with a Geno 2.0 Next Generation test from Helix

Helix, a new genetics start up company in the US, has just announced the launch of the first product on its new pay-as-you-go sequencing platform –  a National Geographic Geno 2.0 Next Generation test. When you order the test the company will sequence your exome. That's the part of your genome which includes all the genes. Your DNA is then stored by the company and you can order additional DNA products as and when they become available. These will include reports on nutrition, health and fitness produced by other partner companies. Presumably Helix hopes that customers will be encouraged to pay for enough add-on products to recoup the costs of the exome sequencing. At present the Helix test is only available to US residents.
Helix is using a technology called Exome+ which they describe as follows:
The “exome” is comprised of all the DNA that encodes for protein—and because proteins are the machinery of your cells, the exome represents some of the most important and well-studied pieces of your DNA. But the exome is only part of your DNA story. The genetic experts at Helix have identified other important information-rich areas to sequence (hence, Exome+).
However, rather than using all the data from the exome, the Helix Geno Next Generation test is done using just a subset of these SNPs. Helix explain that they "provide National Geographic with more than 200,000 markers from your autosomal chromosomes, the Y-chromosome, and mitochondrial DNA". The breakdown of the markers tested is provided on the product page:
  • Maternal line: over 3,000 markers on mitochondrial DNA
  • Paternal line (for males): over 10,000 markers on the Y chromosome
  • Hominin and regional: over 200,000 markers across the entire genome
It is not clear how many of the SNPs used by Helix overlap with the SNPs used on the current Geno 2.0 NextGen test from the Genographic Project. The chip used for the standard Geno 2.0 NextGen test has around 700,000 autosomal SNPs, 20,000 Y-SNPs and 4000 mtDNA SNPs. The new Helix test therefore provides less coverage than the existing test. This is presumably because there are fewer ancestry informative markers in the exome.

Customers in the US who wish to order a Genographic test now have no other option but to buy the Geno Next Generation Helix Kit. The higher-resolution Geno 2.0 Next Gen test is still available for customers outside the US. Presumably the company will wait and see how the test fares in the US before deciding whether or not to roll it out to the rest of the world.

Unfortunately because the new Helix test covers so few markers it will no longer be possible for US customers to transfer their results to the Family Tree DNA Family Finder database to search for genealogical matches. They will also not be able to upload their results to the free third-party websites such as GedMatch and DNA.Land.

At the moment Helix customers cannot access their raw data. The website says that they are actively working on a feature to allow customers to "purchase access to the raw data set containing your complete DNA sequence data in 2017". It remains to be seen how much this will cost, but if someone is interested in having their exome sequenced then the Helix test might turn out to be a cost-effective way of doing so.

The concept of pay as you go sequencing is interesting but I would have thought it would make sense to wait until the full information is available from whole genome sequencing rather than ordering an exome sequencing test. In the current Full Genomes Corporation sale it is already possible to buy 30x whole genome sequencing for $1250, and 15x whole genome sequencing for $795. The Full Genomes test includes a full interpretation of the Y-chromosome data. The raw autosomal data can be uploaded to Promethease for a small fee of $5 for health reports. Veritas Genetics offers a 30x whole genome sequencing test for $999 which includes health and trait reports. No Y-chromosome interpretation is provided though this can be purchased though YFull for $49.  However, the Veritas test needs to be authorised by a doctor. No doubt the cost of whole genome sequencing will come down in price in the next few years to a more affordable level.

Update 1st February 2017
Further information about the new Helix test is provided in an article on GenomeWeb Helix readying for summer launch of genomics apps addressing broad consumer interests (31 January 2017). Here are two relevant quotes from the article:

"In November Helix began offering National Geographics' ancestry product Geno 2.0. Customers can order the spit kit through Helix and view their results online, although not yet through Helix's platform. This is allowing Helix to test out its ability to handle a large product launch, so by summer it can smoothly introduce multiple apps across all six categories."

"The company [Helix] has projected that the initial sequencing and app-based interpretation will cost around $200, a price point that Helix believes more healthy people — those who otherwise don't have a medical reason to seek more expensive testing — will choose to pay out of pocket. Helix is also betting that its app-based model will uncover novel uses for genomic data and enable a responsible path to delivering this information to consumers, by placing them in charge of what they want to learn: something "fun" (and some might argue light on scientific validity), like their wine tasting profile, or something that can impact their future wellbeing and that of their families, such as their hereditary risk for breast cancer."

Update 2nd December 2017
For further details about the Helix platform see the Helix Personal Genomics Platform White Paper.

According to this blog post from Razib Khan it is expected that the raw data download will be available early in 2018 and will cost in the region of $600.

A further update from Razib Khan explains that the test will sequence around 30,000,000 markers.

Further reading
With thanks to Gerard Corcoran, James Kane, David Mittelman and Ann Turner. 

Tuesday, 20 January 2015

What is the current size of the consumer genomics market?

The subject of how many people have taken a DNA test is always the source of much speculation, and reliable figures are hard to come by. However, in a report published this week by GenomeWeb Spencer Wells, director of National Geographic's Genographic Project anticipates that "the 3 millionth person" [will] test him or herself during the next few months". In the same article Roberta Estes, who writes the popular DNAeXplained blog, suggests that the three million milestone might already have been achieved. She notes: "23andMe has stated publicly that it has genotyped 800,000 kits, AncestryDNA and the Genographic Project each has genotyped perhaps more than 700,000, and Family Tree DNA has genotyped close to 120,000 people for its Family Finder autosomal DNA offering alone." I thought I would take a look at the available sources for the different companies to see if it might be possible to verify these figures and provide an estimate of the current total.

23andMe
23andMe state in their media fact sheet that they have genotyped more than 800,000 customers.

The 23andMe test is sold in 56 countries of the world. However, I estimate that about 90% of their customer base is in the US. Canada and the UK are currently the only countries where the 23andMe test includes the health and trait reports.

The Genographic Project
The Genographic Project's home page states, as of today's date, that the project has 705,343 participants.

I understood that the Genographic Project kit could be purchased from any country in the world, but from the dropdown menu in their online shop it would appear that the kit is now sold in just 33 countries.

AncestryDNA
AncestryDNA confirmed in August 2014 that they had tested over 500,000 DNA customers. In a presentation given towards the end of last year by Ken Chahine, Ancestry's senior Vice President and General Manager, he stated that AncestryDNA were selling 30,000 to 50,000 DNA kits per month. If we take the middle figure of 40,000 multiplied by six that gives us a figure of 240,000 kits sold since August 2014, bringing the total up to 740,000.

The AncestryDNA test is currently only sold in America, but there are plans to launch the test in the UK, Ireland, Australia and perhaps other countries later this year.

Family Tree DNA
Family Tree DNA provide details only on the number of different types of tests taken and not the total number of customers. According to their website, as of today's date, their stats are as follows:

- 520,257 Y-chromosome DNA records in the database. The Y-DNA database includes 180,005 people who have tested at least 37 Y-STR markers. The FTDNA database also includes several thousand people who have taken the advanced BIG Y test, a comprehensive Y-chromosome sequencing SNP discovery test. FTDNA almost certainly have the largest Y-chromosome DNA database in the world with samples tested at higher resolution than in any other database.

- 190,105 mitochondrial DNA records in the database. The mtDNA database includes 47,849 people who have taken the full mitochondrial sequence (FMS) test. (This test was previously known as the FGS - full genomic sequence test). FTDNA probably have the world's largest database of full mtDNA genomes.

- The number of autosomal Family Finder tests in the FTDNA database has not been publicly disclosed. It is not clear if the 120,000 figure cited by Roberta Estes in the GenomeWeb article mentioned above is an estimate or an actual figure obtained from FTDNA staff, but the number certainly seems to be in line with my own estimates.

FTDNA sell their tests in theory to any of the 200 or so countries of the world. However, they are unable to ship to Iran and Sudan because of customs restrictions.

FTDNA have partnerships with the European company iGENEA and the Middle Eastern company DNA Ancestry & Family Origin. These partnerships have helped to bring in many non-English-speaking customers from Europe and the Middle East, but again many more who will have tested direct with FTDNA.

iGENEA kit numbers are preceded by the letter E. The iGENEA kit numbers in my mtDNA Haplogroup U4 Project go up to kit no. E17977 so it would appear that nearly 20,000 Europeans have tested through iGENEA. Many Europeans will also have tested directly through FTDNA. (It is in fact considerably cheaper to order direct through FTDNA rather than through iGENEA, but iGENEA do have the advantage of a website which is available in French, German, Spanish and Italian.)

The kits from the Middle East are preceded by the letter M. The highest kit with the M prefix that I can find in the large Arab Tribes DNA Project is kit no. 9658 so there are perhaps around 10,000 people who have tested through the FTDNA affiliate in the Middle East.

Family Tree DNA also have partnerships with a number of smaller companies such as DNA Worldwide and Jewish Voice, though these partnerships probably only account for a few thousand kits. For details on the various prefixes see the ISOGG Wiki article on Family Tree DNA kit numbers.

The international diversity of the FTDNA database can be seen in the huge range of geographical DNA projects, which are run by volunteer project administrators from around the world.

Family Tree DNA are the testing partner for the Genographic Project, and all the Geno 2.0 tests are processed in FTDNA's lab in Houston, Texas. Genographic Project participants have the option of transferring their results into the FTDNA database. Genographic Project kit numbers are preceded by the letter N. The highest Genographic Project kit number in the Haplogroup U4 Project is kit number N129937. We therefore know that around 130,000 Genographic Project customers have transferred their results to FTDNA.

Family Tree DNA are the only company who will accept autosomal transfers from other testing companies. They can accept transfers for people who have tested at both AncestryDNA and 23andMe. However, 23andMe transfers can only be accepted if the test was done on the version 3 chip which was sold between November 2011 and November 2013. Kit numbers for the autosomal transfers are prefixed by the letter B. The same prefix is also used for Y-DNA transfers from AncestryDNA and DNA Heritage. AncestryDNA no longer offer Y-STR testing. FTDNA purchased the British company DNA Heritage in April 2011. The highest B kit I can find in my projects is B39616 in the Haplogroup U4 Project, so it would appear that there are getting on for 40,000 third-party transfers in the FTDNA database. Both DNA Heritage and AncestryDNA only ever had quite small Y-DNA databases, and in any case not everyone transferred their Y-DNA results, so I would guess that the majority of the third-party transfers (perhaps in the region of 35,000) are autosomal results from 23andMe and AncestryDNA. It is not clear if the third-party transfers are included in the estimate of the size of the FTDNA Family Finder database or if these transfers are in addition to the autosomal tests processed directly by FTDNA.

It is impossible from these figures to determine precisely how many individuals there are in the Family Tree DNA database because many people who have ordered a Y-DNA test will also have gone on to order a Family Finder test and/or a mitochondrial DNA test and vice versa. The kit numbers probably provide the closest approximation of the number of people in the database. My highest FTDNA kit number is kit number 394825 in the Devon DNA Project. It may well be that the 400,000 milestone has already been passed. If we assume that there are 400,000 FTDNA kits, 130,000 Genographic transfers, 20,000 iGENEA kits, 10,000 kits from FTDNA's Middle Eastern partner, and 5,000 miscellaneous kits, we get a figure of 565,000 which is probably a reasonable estimate of the number of individuals in the FTDNA database.

Other companies
In addition to the big four companies there are a number of other smaller companies such as BritainsDNA, Oxford Ancestors and GeneBase which sell genetic ancestry tests direct to the consumer. A full list of DNA testing companies can be found in the ISOGG Wiki. However, none of these smaller companies disclose the size of their databases, and many of the people who've tested with the smaller companies have retested with one of the big four companies. I hesitate to estimate the number of people tested with these different companies but I do not think the figure can be more than 50,000 and is very likely to be much less than this.

What is the total?
To sum up, the total number of individuals tested at each of the four big companies is as follows;

Genographic Project  705,343
23andMe                    800,000+
Family Tree DNA      565,000 (DK estimate)
AncestryDNA            740,000 (DK estimate)

If we add all these figures together we get a total of 2,810,343. However, this figures makes no allowance for the significant overlap in the four databases as there are many people who have tested at multiple companies. For example, I've had my own DNA tested at 23andMe, Family Tree DNA and AncestryDNA. We can subtract the 130,000 people who have transferred their Genographic results to FTDNA and we can perhaps estimate that about 35,000 people have transferred autosomal DNA results to FTDNA.  That brings the total down to 2,645,343. There is probably more overlap than I've allowed for, but it does seem very likely that there are currently around two and a half million people in the world who have paid for a DNA test with the big four companies. It will be interesting to see what these figures look like this time next year.

© 2015 Debbie Kennett

Friday, 25 April 2014

The new 2014 Y-DNA haplotree has arrived!

Today saw the launch of Family Tree DNA's new 2014 Y-DNA haplotree which has been created in partnership with National Geographic's Genographic Project. If you've tested with Family Tree DNA you will find the new tree by going to your personal page and clicking on "haplotree and SNPs". Below is a screenshot of the upper portion of the tree for haplogroup R:


The tree is too large to fit into a single screenshot. Here is the relevant portion of the tree for my dad who is R-Z12, a sub-branch of U106.


Note that this is very much an interim tree. It is based on SNPs tested with the Genographic Project's Geno 2.0 chip, and the cut off date for inclusion of SNPs is November 2013. The new tree does not include all the thousands of new SNPs identified from testing with Big Y, Full Genomes and Chromo 2. The tree will eventually be much more comprehensive but FTDNA are being careful about the data they use from other sources and are insisting that SNPs are only added from published data and raw data that they have personally verified rather than from interpreted data. They have promised that at least one update will be released this year. The FTDNA Learning Center will eventually be updated with information about the new haplotree. If you have questions about a particular SNP that is in the wrong place on the tree or if you spot any other errors you should send an e-mail to the FTDNA help desk with Y-Tree in the subject line.

FTDNA are now recommending SNPs for people to test. I've only had a chance to look briefly at the SNP recommendations for a few project members. It is apparent that in some cases the SNPs that are recommended for testing are not appropriate. SNPs are only recommended if they pass certain percentage thresholds and there might well be a more appropriate downstream SNP that would be more suitable. If you are interested in ordering single SNP testing, make sure you join the appropriate Y-DNA haplogroup project and seek advice from the project administrators. If not, you could end up wasting money ordering unnecessary SNPs.

The following information has been provided by Family Tree DNA.

 FAST FACTS
• Created in partnership with National Geographic’s Genographic Project
• Used GenoChip containing ~10,000 previously unclassified Y-SNPs
• Some of those SNPs came from Walk Through the Y and the 1000 Genome Project
• Used first 50,000 high-quality male Geno 2.0 samples
• Verified positions from 2010 YCC by Sanger sequencing additional anonymous samples
• Filled in data on rare haplogroups using later Geno 2.0 samples

Statistics
• Expanded from approximately 400 to over 1200 terminal branches
• Increased from around 850 SNPs to over 6200 SNPs
• Cut-off date for inclusion for most haplogroups was November 2013

Total number of SNPs broken down by haplogroup:
A 406
B 69
BT 8
C 371
CT 64
D 208
DE 16
E 1028
F 90
G 401
H 18
I 455
IJ 29
IJK 2
J 707
K 11
K(xLT) 1
L 129
LT 12
M 17
N 168
NO 16
O 936
P 81
Q 198
R 724
S 5
T 148

myFTDNA Interface
• Existing customers receive free update to predictions and confirmed branches based on existing SNP test results.
• Haplogroup badge updated if new terminal branch is available
• Updated haplotree design displays new SNPs and branches for your haplogroup
• Branch names now listed in shorthand using terminal SNPs
• For SNPs with more than one name, in most cases the original name for SNP was used, with synonymous SNPs listed when you click "More…"
• No longer using SNP names with .1, .2, .3 suffixes. Back-end programming will place SNP in correct haplogroup using available data.
• SNPs recommended for additional testing are pre-populated in the cart for your convenience. Just click to remove those you don’t want to test.
• SNPs recommended for additional testing are based on 37-marker haplogroup origins data where possible, 25- or 12-marker data where 37 markers weren't available.
• Once you've tested additional SNPs, that information will be used to automatically recommend additional SNPs for you if they’re available.
• If you remove those prepopulated SNPs from the cart, but want to re-add them, just refresh your page or close the page and return.
• Only one SNP per branch can be ordered at one time – synonymous SNPs can possibly [be] ordered from the Advanced Orders section on the Upgrade Order page.
• Tests taken have moved to the bottom of the haplogroup page.

Coming attractions
• Group Administrator Pages will have longhand removed.
• At least one update to the tree to be released this year.
• Update will include: data from Big Y, relevant publications, other companies' tests from raw data.
• We'll set up a system for those who have tested with other big data companies to contribute their raw data file to future versions of the tree.
• We're committed to releasing at least one update per year.
• The Genographic Project is currently integrating the new data into their system and will announce on their website when the process is complete in the coming weeks. At that time, all Geno 2.0 participants’ results will be updated accordingly and accessible via the Genographic Project website.

BACKGROUND
Family Tree DNA created the 2014 Y-DNA Haplotree in partnership with the National Geographic Genographic Project using the proprietary GenoChip. Launched publicly in late 2012, the chip tests approximately 10,000 Y-DNA SNPs that had not, at the time, been phylogenetically classified.

The team used the first 50,000 male samples with the highest quality results to determine SNP positions. Using only tests with the highest possible “call rate” meant more available data, since those samples had the highest percentage of SNPs that produced results, or “calls.”

In some cases, SNPs that were on the 2010 Y-DNA Haplotree didn’t work well on the GenoChip, so the team used Sanger sequencing on anonymous samples to test those SNPs and to confirm ambiguous locations.

For example, if it wasn’t clear if a clade was a brother (parallel) clade, or a downstream clade, they tested for it.

The scope of the project did not include going farther than SNPs currently on the GenoChip in order to base the tree on the most data available at the time, with the cutoff for inclusion being about November of 2013.

Where data were clearly missing or underrepresented, the team curated additional data from the chip where it was available in later samples. For example, there were very few Haplogroup M samples in the original dataset of 50,000, so to ensure coverage, the team went through eligible Geno 2.0 samples submitted after November, 2013, to pull additional Haplogroup M data. That additional research was not necessary on, for example, the robust Haplogroup R dataset, for which they had a significant number of samples.

Family Tree DNA, again in partnership with the Genographic Project, is committed to releasing at least one update to the tree this year. The next iteration will be more comprehensive, including data from external sources such as known Sanger data, Big Y testing, and publications. If the team gets direct access to raw data from other large companies’ tests, then that information will be included as well. We are also committed to at least one update per year in the future.

Known SNPs will not intentionally be renamed. Their original names will be used since they represent the original discoverers of the SNP. If there are two names, one will be chosen to be displayed and the additional name will be available in the additional data, but the team is taking care not to make synonymous SNPs seems as if they are two separate SNPs. Some examples of that may exist initially, but as more SNPs are vetted, and as the team learns more, those examples will be removed.

In addition, positions or markers within STRs, as they are discovered, or large insertion/deletion events inside homopolymers, potentially may also be curated from additional data because the event cannot accurately be proven. A homopolymer is a sequence of identical bases, such as AAAAAAAAA or TTTTTTTTT. In such cases it’s impossible to tell which of the bases the insertion is, or if/where one was deleted. With technology such as Next Generation Sequencing, trying to get SNPs in regions such as STRs or homopolymers doesn’t make sense because we’re discovering non-ambiguous SNPs that define the same branches, so we can use the non-ambiguous SNPs instead.  Some SNPs from the 2010 tree have been intentionally removed. In some cases, those were SNPs for which the team never saw a positive result, so while it may be a legitimate SNP, even haplogroup defining, it was outside of the current scope of the tree. In other cases, the SNP was found in so many locations that it could cause the orientation of the tree to be drawn in more than one way. If the SNP could legitimately be positioned in more than one haplogroup, the team deemed that SNP to not be haplogroup defining, but rather a high polymorphic location.

To that end, SNPs no longer have .1, .2, or .3 designations. For example, J-L147.1 is simply J-L147, and I-147.2 is simply I-147.  Those SNPs are positioned in the same place, but back-end programming will assign the appropriate haplogroup using other available information such as additional SNPs tested or haplogroup origins listed. If other SNPs have been tested and can unambiguously prove the location of the multi-locus SNP for the sample, then that data is used. If not, matching haplogroup origin information is used.

We will also move to shorthand haplogroup designations exclusively. Since we’re committing to at least one iteration of the tree per year, using longhand that could change with each update would be too confusing.  For example, Haplogroup O used to have three branches: O1, O2, and O3. A SNP was discovered that combined O1 and O2, so they became O1a and O1b.

There are over 1200 branches on the 2014 Y Haplogroup tree, as compared to about 400 on the 2010 tree. Those branches contain over 6200 SNPs, so we’ve chosen to display select SNPs as “active” with an adjacent “More” button to show the synonymous SNPs if you choose.

The Genographic Project is currently integrating the new data into their system and will announce on their website when the process is complete in the coming weeks.  At that time, all Geno 2.0 participants’ results will be updated accordingly and will be accessible via the Genographic Project website.

QUOTES
Elliot Greenspan has provided the following quotes in conversation with Janine Cloud, Family Tree DNA's GAP Liaison and Events Co-ordinator:

"I want it to be the most accurate tree it can be, but I also want it to be interesting. That's the key. Historical relevance is what we're to discover. Anthropological relevance. It's not just who has the largest tree, it's who can make the most sense out of what you have [that] is important."

"This year we're committing to launching another tree. This tree will be more comprehensive, utilizing data from external sources: known Sanger data, as well as data such as Big Y, and if we have direct access to the raw data to make the proof (from large companies, such as the Chromo2) or a publication, or something of that nature. That is our intention that it be added into the data."

"We’re definitely committed to update at least once per year. Our intention is to use data from other sources, as well as any SNPs we can, but it must be well-vetted. NGS and SNP technology inherently has errors. You must curate for those errors otherwise you’re just putting slop out to customers. There are some SNPs that may bind to the X chromosome that you didn’t know. There are some low coverages that you didn’t know."

"With technology such as this [next-generation sequencing] you're able to overcome the urge to test only what you’re likely to be positive for, and instead use the shotgun method and test everything. This allows us to make the discovery that SNPs are not nearly as stable as we thought, and they have a larger potential use in that sense."

"Not only does the raw data need to be vetted but it needs to make sense. Using Geno 2.0, I only accepted samples that had the highest call rate, not just because it was the best quality but because it was the most data. I don't want to be looking at data where I'm missing potential information A, or I may become confused by potential information B. That is something that will bog us down. When you’re looking at large data sets, I’d much rather throw out 20% of them because they’re going to take 90% of the time than to do my best to get one extra SNP on the tree or one extra branch modified, that is not worth all of our time and effort. What is, is figuring out what the broader scope of people are, because that is how you break down origins. Figuring one single branch for one group of three people is not truly interesting until it's 50 people, because 50 people is a population. Three people may be a family unit. You have to have enough people to determine relevance. That's why using large datasets and using complete datasets are very, very important."

Update 27 April 2014
A recording of the Family Tree DNA webinar presented by Elise Friedman on the launch of the new 2014 haplotree is now available online and can be accessed here (free registration required).

Related blog posts
- A confusion of SNPs

Wednesday, 12 December 2012

Genographic results from the UK

The first results from Geno 2.0, the new DNA test from the Genographic Project, are now starting to appear. A genetic genealogy friend in the UK has very kindly agreed to share screenshots of his results with me for publication on this blog. One of his parents is English and the other is from the Philippines so he has some very interesting results. Each participant is a given a very cool infographic summarising their results which they can share with their friends.
These are the pages which tell the personal genetic story of the participant.

The two reference populations with which this participant most closely matches are Vietnam and Romania. These seem rather odd selections and don't match his documented ancestry from England and the Philippines, but perhaps there are insufficient reference populations in the database to give accurate matches. 
This close up provides details of the British reference population used by the Genographic Project.
A fun part of the test is that you are told your percentages of Neanderthal and Denisovan ancestry.
For the Y-DNA results you get a nice map showing the migratory path of the different branches of the Y-DNA tree. This is the map showing the path of  U106, one of the major branches of the R1b tree.
We can then follow the journey of U198, one of the subclades of U106.
This rather nice heat map shows the distribution of U198, which appears to be found almost exclusively in the British Isles and north-western France. It would be helpful to have the references that were used to compile the map. Perhaps that information will be added later.
For the mitochondrial DNA there is a description of the haplogroup, which in this case is haplogroup F, reflecting the participant's maternal ancestry from the Philippines.
There is a map showing the migratory path of haplogroup F. 
 There is also a heat map showing the places where haplogroup F is mostly found, though again it would be useful to have a list of the sources used.
Genographic results can be transferred free of charge to the Family Tree DNA database. CeCe Moore has blogged about her own results and has also included detailed instructions on the process of transferring results to FTDNA. We will no doubt learn much more as people test and contribute their results to research. Genographic results will be updated on a regular basis as more results are received and more reference populations are added to the database. For further information on the Genographic Project visit the Genographic website.

Websites

Thursday, 22 November 2012

First results from Geno 2.0

A few people have started to receive their first results for the new Geno 2.0 test from the Genographic Project. CeCe Moore has posted some screenshots on her blog showing results for the autosomal DNA component of the test. Dave Dowell has blogged about his mtDNA results and has published an example of a "heat map" for haplogroup H. We are still waiting to see the first Y-DNA results.

I have not ordered the Geno 2.0 test for myself as I have already taken the full mitochondrial sequence test with Family Tree DNA.  The ethnicity percentages from the autosomal part of the test will not yield any meaningful information and should only confirm that I am "British", which I already know from my family history research! As a female I do not have a Y-chromosome and I would, therefore, not be able to receive any Y-DNA results. However, if you are interested in your deep ancestry and wish to know your mtDNA and Y-DNA haplogroups then the Geno 2.0 test is a good choice. The Y-DNA and mtDNA results can also be transferred to Family Tree DNA where you can join the relevant haplogroup projects and order additional testing for genealogical purposes. The Geno 2.0 test will provide very detailed Y-DNA haplogroup assignments and has essentially replaced the old deep clade test offered by Family Tree DNA.

I contributed my mtDNA results from Family Tree DNA to the first phase of the Genographic Project. I am still able to log into my Genographic account to access my results. The website has been updated and I am now getting the same screenshot as the Geno 2.0 participants. However, I note that I have been downgraded to a simple haplogroup U rather than a U4, and I am no longer able to access any information on haplogroup U4. It may be that the new haplogroup pages have not yet gone live and my haplogroup will be adjusted when this has been done.



© 2012 Debbie Kennett


Sunday, 4 November 2012

ASHG abstracts

The annual meeting of the American Society of Human Genetics will take place from 6 - 10 November in San Francisco. The posters can be searched online from the ASHG meeting website. The following three abstracts will be of particular interest to the genetic genealogy community.

The GenoChip: a new tool for genetic anthropology
S. Wells, E. Greenspan, S. Staats, T. Krahn, C. Tyler-Smith, Y. Xue, S. Tofanelli, P. Francalacci, F. Cucca, L. Pagani, L. Jin, H. Li, T. G. Schurr, J. B. Gaieski, C. Melendez, M. G. Vilar, A. C. Owings, R. Gomez, R. Fujita, F. Santos, D. Comas, O. Balanovsky, E. Balanovska, P. Zalloua, H. Soodyall, R. Pitchappan, G. Arun Kumar, M. F. Hammer, B. Greenspan, E. Elhaik

 Background: The Genographic Project is an international effort aimed at charting human history using genetic data. The project is non-profit and non-medical, and through the sale of its public participation kits it supports cultural preservation efforts in indigenous and traditional communities. To extend our knowledge of the human journey, interbreeding with ancient hominins, and modern human demographic history, we designed a genotyping chip optimized for genetic anthropology research. Methods: Our goal was to design, produce, and validate a SNP array dedicated to genetic anthropology. The GenoChip is an Illumina HD iSelect genotyping bead array with over 130,000 highly informative autosomal and X-chromosomal SNPs ascertained from over 450 worldwide populations, ~13,000 Y-chromosomal SNPs, and ~3,000 mtDNA SNPs. To determine the extent of gene flow from archaic hominins to modern humans, we included over 25,000 SNPs from candidate regions of interbreeding between extinct hominins (Neanderthal and Denisovan) and modern humans. To avoid any inadvertent medical testing we filtered out all SNPs that have known or suspected health or functional associations. We validated the chip by genotyping over 1,000 samples from 1000 Genomes, Family Tree DNA, and Genographic Project populations. Results: The concordance between the GenoChip and the 1000 Genomes data was over 99.5%. The GenoChip has a SNP density of approximately (1/100,000) bases over 92% of the human genome and is highly compatible with Illumina and Affymetrix commercial platforms. The ~10,000 novel Y SNPs included on the chip have greatly refined our understanding of the Y-chromosome phylogenetic tree. By including Y and mtDNA SNPs on an unprecedented scale, the GenoChip is able to delineate extremely detailed human migratory paths. The autosomal and X-chromosomal markers included on the GenoChip have revealed novel patterns of ancestry that shed a detailed new light on human history. Interbreeding analysis with extinct hominids confirmed some previous reports and allowed us to describe the modern geographical distribution of these markers in detail. Conclusions: The GenoChip is the first genotyping chip completely dedicated to genetic anthropology with no known medically relevant markers. We anticipate that the large-scale application of the GenoChip using the Genographic Project’s diverse sample collection will provide new insights into genetic anthropology and human history.
View source

People of the British Isles: An analysis of the genetic contributions of European populations to a UK control population
S. Leslie, B. Winney, G. Hellenthal, S. Myers, P. Donnelly, W. Bodmer

There is much interest in fine scale population structure in the UK, as a signature of historical migration events and because of the effect population structure may have on disease association studies. Population structure appears to have a minor impact on the current generation of genome-wide association studies, but will probably be important for the next generation of studies seeking associations to rare variants. Furthermore there is great interest in understanding where the British people came from. Thus far genetic studies have been limited to a small number of markers or to samples not collected to specifically address these questions. A natural method for understanding population structure is to control and document carefully the provenance of samples. We describe the collection of a cohort of rural UK samples (The People of the British Isles), aimed at providing a well-characterised UK control population. This will be a resource for research community as well as providing fine-scale genetic information on the history of the British. Using a novel clustering algorithm, approximately 2000 samples were clustered purely as a function of genetic similarity, without reference to their known sampling locations. When each individual is plotted on a UK map, there is a striking association between inferred clusters and geography, reflecting to a major extent the known history of the British peoples. A similar analysis is performed on samples from different parts of Europe. Using the European samples as ‘source populations’ we apply a novel algorithm to determine the proportion of the genomes within each of the derived British clusters that are most closely related to each of the source populations. Thus we can observe the relative contribution (under our model) of each of these European populations to the genomes of samples in different regions of Britain. Our results strikingly reflect much of the known historical and archaeological record while raising some important questions and perhaps answering others. We believe this is the first detailed analysis of very fine-scale genetic structure and its origin in a population of very similar humans. This has been achieved through both a careful sampling strategy and an approach to analysis that accounts for linkage disequilibrium.
View source

Inferring Y Chromosome Phylogeny by Sequencing Diverse Populations
G. D. Poznik, P. A. Underhill, B. M. Henn, M. C. Yee, E. Sliwerska, G. M. Euskirchen, L. Quintana-Murci, E. Patin, M. Snyder, J. M. Kidd, C. D. Bustamante

The male-specific region of the Y chromosome (MSY) harbors the longest stretch of non-recombining DNA in the human genome and is therefore a unique tool that enables the tracking of migrations and inference of demographic history. We have sequenced 69 male samples from nine globally diverse populations, including three African hunter-gatherer groups. Due to inefficient selection, a relatively high mutation rate, and a small effective population size, the Y chromosome is particularly subject to drift. It has accumulated large expanses of highly repetitive sequence, which pose considerable challenge within a short read sequencing paradigm. To overcome this hurdle, we have built an informatics pipeline to reliably call Y chromosome alleles from moderate coverage short read shotgun sequence data. First, we defined a callability mask, learned from the mapping quality and depth of coverage patterns in the data, and then we tuned base-pair level quality control thresholds. Based on 13,000 provisional SNP calls, we inferred a tree of the 69 sequenced Y chromosomes. Using this tree, we then called individual genotypes for each SNP with a custom-built, phylogeny-aware, EM algorithm. With these high quality calls in hand, samples were assigned haplogroup labels using standard YCC nomenclature; 29 distinct named haplogroups were represented. We find that the maximum likelihood tree we construct recapitulates the extant Y chromosome phylogeny, thus confirming the fruits of decades of work based on ascertained SNPs. Further, we resolve a major long-standing polytomy by identifying a variant for which one haplogroup retains the ancestral allele, whereas its brother clades share the derived allele, thus indicating common ancestry and uniting the latter two branches. This finding has been confirmed by genotyping a larger panel. Finally, we estimate the MSY rate of mutation recurrence and the time to the most recent common ancestor of the sampled chromosomes.
View source

Saturday, 18 August 2012

Geno 2.0 update

I wrote about the launch of the new Geno 2.0 test from the Genographic Project at the end of July. Some further details about the test have become available in the last few weeks. The genetic genealogy bloggers, CeCe Moore and Roberta Estes, have both received e-mails from Spencer Wells, National Geographic's Explorer-in-Residence, with further details on the SNPs (markers) to be used in the project and the degree of community involvement. Links to the posts are provided below:

- CeCe Moore   A short update from Spencer Wells on Geno 2.0 (30 July 2012)
- Roberta Estes  Geno 2.0 answers from Spencer Wells (30 July 2012)
- CeCe Moore   More information from Spencer Wells on Geno 2.0 (31 July 2012)
- Roberta Estes  Geno 2.0, WTY, mtDNA full sequence participants, and more (31 July 2012)

Charles Moore, the administrator of the haplogroup R1b-U106 project, and his co-administrator Mike Maddi had the opportunity to have lunch with Bennett Greenspan, the CEO of Family Tree DNA, a couple of weeks ago. Charles reported back to the U106 mailing list on 6th August and he has very kindly given me permission to reprint his posting here which provides some additional technical information about the new Geno 2.0 test.
R1b-U106 Project Co-Administrator Mike Maddi and I had lunch with Bennett Greenspan on Saturday. Bennett gave us a tour of the renovated lab, including the new DNA storage machine with windows that allowed us to watch it perform its various robotic tasks, efficiently filling little wells in plates with lots of wonderful little DNA samplings.

The discussion naturally moved quickly towards the new National Geographic "Genographic Project" Geno 2.0 test.

Bennett said that yes, individual SNP testing will remain at FTDNA, and when significant or terminal branch SNPs are discovered via Geno 2.0 that are not currently testable on an individual basis at FTDNA, they would be made testable. Since we do not yet know the number of SNPs we are talking about, it's not yet feasible to estimate when these will become available, but likely around the end of the year.

But testers interested in testing lots of SNPs, for whatever reasons, should sign up for the National Geographic Geno 2.0 test. Going forward, this test will be the method for accomplishing this objective.

Bennett added that lots of SNPs from Asian labs, and Near Eastern/Mediterranean labs, that we are mostly not otherwise familiar with, are included on the chip. Additionally about 5,000 newly identified SNPs from the 1000 Genomes project have been added to the chip. And of course, the chip tests mtDNA and autosomal DNA as well as Y DNA. Aside from the reports about 12,000 Y SNPs on the chip, Bennett added that about 1,000 of them are already known to be below Haplo R1, however many are likely synonymous with current SNPs on the tree.

Bennett did say that the POSITIVE results from Y SNPs on the Geno 2.0 test, may be re-merged with one's existing FTDNA account, and thereby will also show up on the Project's public SNP list. As Administrators, Mike and I were very grateful for this answer!

23 public WTY testers' samples, and approximately 300 FMS (aka mtDNA FGS) testers' samples were used to help verify the chip. These testers will be notified in the next few weeks, and they will receive Geno 2.0 refunds if they have already ordered Geno 2.0, Bennett confirmed. They will also receive Geno 2.0 accounts that can be re-merged into their FTDNA accounts as well.
WTY is an abbreviation for Walk Through the Y, a SNP discovery project at Family Tree DNA. Further information about WTY can be found on the Walk Through the Y page in the ISOGG Wiki. FMS is an abbreviation for the full mitochondrial sequence test which is also known as the full genomic sequence (FGS) test. This test is available from Family Tree DNA and it provides a reading of all 16,569 bases in the mitochondrial genome.

I previously reported that the Geno 2.0 kits would be available for order from Family Tree DNA. It appears that this is not the case and I have now updated my previous blog post accordingly.

Former Genographic Project participants can receive a discount of $30 off the cost of the new kit for a limited period. The discount can only be obtaining by calling the customer service line so this is probably only a realistic option for participants living in North America.

Geno 2.0 can be pre-ordered via the Genographic Project website.

There is also information about the test on the Family Tree DNA website.

A summary of the key features of the new Geno 2.0 test can be found on the Genographic Project page in the ISOGG Wiki.

© Debbie Kennett 2012

Thursday, 26 July 2012

Ancestry SNPs galore

Today sees the launch of  Geno 2.0, an exciting new DNA test which marks phase two of the Genographic Project, a scientific research project run by National Geographic in partnership with Family Tree DNA and IBM. The launch has been covered by a number of bloggers in America who attended the pre-launch presentation. They each cover the test from slightly different perspectives and all the posts are well worth reading.

- Roberta Estes     National Geographic Geno 2 announcement - the human story

- Roberta Estes     Geno 2.0 - Q&A with Bennett Greenspan
- CeCe Moore        National Geographic and Family Tree DNA announce Geno 2.0 

- Blaine Bettinger   National Geographic and Family Tree DNA announce Geno 2.0
- Judy Russell       Geno 2.0 launches 
- Emily Aulicino     National Geographic announces new DNA test
- Razib Khan        The Genographic Project: on to the autosome!

The new Geno 2.0 test is a deep ancestry test. It complements the existing DNA tests that are traditionally used by genealogists and does not replace them. The new test is looking at special markers known as SNPs (pronounced 'snips'). These markers can tell us how much of our DNA we share in common with other populations from around the world. Both males and females will discover their mitochondrial DNA haplogroup. Males will discover their Y-DNA haplogroup. Haplogroups represent branches of the human family tree. The inclusion of new SNPs in the Geno 2.0 test will allow scientists to define not just the branches of the family tree, but also the twigs (subclades) which make up those branches.

For surname projects we use a different marker known as a Y-STR marker. Y-STR markers are like the leaves on a tree and we use Y-STR tests to place the leaves on the tree by grouping matching results in genetic families. If you are a male and are interested in using a DNA test to help with your genealogy research you should take a Y-STR test with one of the many surname or geographical projects at Family Tree DNA.

The new Geno 2.0 test will, however, effectively serve as a replacement for the old Y-DNA deep clade test from Family Tree DNA which provided detailed subclade assignments. The old deep clade test looked at a small handful of SNPs and it was necessary to order new SNPs à la carte as and when they were discovered. For just a little bit more money the new Geno 2.0 test covers 12,000 Y-DNA SNPs all in one go, and includes many new SNPs that were not previously available. It is set to transform our knowledge of the Y-DNA haplogroups.

The Geno 2.0 test covers 3,200 mitochondrial SNPs (about 19% of the mitochondrial genome). If you are only interested in learning your mtDNA haplogroup assignment then you should take the Geno 2.0 test. If you are interested in using mtDNA for matches within a genealogical timeframe (within the last 400 years) or if you are interested in participating in mtDNA scientific research then you will need to order a full mitochondrial sequence (FMS) test from Family Tree DNA. The FMS test sequences the entire mitochondrial genome (100%) which consists of 16,569 bases.

It will be interesting to see how the new Genographic Project test impacts on the other SNP tests that are currently available. Autosomal SNP testing is currently offered by 23andMe, Family Tree DNA and Ancestry.com. Their tests look at many more SNP markers than the new Geno 2.0 test. 23andMe tests one million SNPs. Family Finder and Ancestry test around 700,000 SNPs, I have written previously about my experiences with the 23andMe test and the Family Finder test. Both these tests can be used to find matches with genetic cousins within the last five generations or so. The 23andMe test also provides information on health issues. Ancestry.com have recently launched their own autosomal DNA test. I am currently waiting to receive my results, and will be reporting on this test in due course. The Geno 2.0 test cannot be used for cousin matching as it does not test enough SNPs to make confident relationship predictions so these tests will continue to be useful for genealogical purposes.

The new Geno 2.0 test will, however, be the test of choice for anyone interested in learning about their ethnicity. The existing autosomal DNA tests only provide very vague information on ethnicity giving percentages of Asian, European and African admixture. The Genographic Project, with the power of its huge database (524,000 tests from 140 countries) and the focus on ancestry informative markers, should in theory be able to provide highly detailed ethnicity breakdowns and they will effectively wipe out all the competition. This new test might hit Ancestry particularly hard. Ancestry have invested a huge amount of money in their new autosomal DNA test and have specifically targeted their test at the large American market, where there is a lot of interest in ethnicity testing. Despite the large investment the early reports suggest that there are many problems with the Ancestry.com tests which have yet to be ironed out.

The British company BritainsDNA offers a deep ancestry SNP test which looks at just 200 Y-DNA and 200 mtDNA SNPs for £200 (no autosomal SNPs are included). In comparison the Geno 2.0 test covers 143,000 SNPs (including 12,000 Y-DNA SNPs and 3,200 mtDNA SNPs) and costs just $199.95  (£128). It is difficult to see how BritainsDNA can possibly compete with the Genographic Project.

Further details about the Geno 2.0 test can be found on the Genographic Project website. Pre-orders are now being accepted. If you try to place an order you are told that the kit will be shipped on 30th October 2012.


See also my blog post dated 18th August which provides additional information on Geno 2.0.  

© 2012 Debbie Kennett

Monday, 21 December 2009

A lecture by Dr Spencer Wells at the National Geographic Store in London

On Sunday 13th December I was privileged to attend a lecture by Dr Spencer Wells at the National Geographic Store in London. Spencer Wells is a National Geographic Explorer-in-Residence and the Director of the Genographic Project, an exciting five-year scientific research programme which is attempting to compile the evolutionary human family tree by collecting DNA samples from around the world. The historical information in our DNA can also tell us about the migratory journeys of our ancient ancestors who left Africa some 60,000 years ago. This brief video provides an introduction to the Project.

The Genographic Project was launched in April 2005. To date over 60,000 DNA samples have been collected from indigenous populations around the world. The general public are also encouraged to take part in the project by purchasing a public participation kit. The response has been overwhelming, and has exceeded all expectations. Over 10,000 kits were sold on the very first day! Today over 330,000 public participation kits have been sold in 130 different countries. The research team have only just started to mine the data from the public kits, and scientific papers are promised in due course. All the data from the project will eventually be made public.

The Legacy Fund is an important component of the project. Proceeds from the sales of the kits are used to fund further field research and to support indigenous conservation and revitalisation projects. Dr Wells showed us some examples of the type of projects supported. In Sierra Leone funds have been used to document the oral poetry of the indigenous population. In South America work is under way to catalogue the native plants and their traditional uses. In Australia work is being done to record and archive traditional music.

Dr Wells gave us a fascinating insight into the difficulties of collecting samples from some of the more remote countries in the world. Many of the countries visited have been off limits to outside researchers for a long time because of civil war or rebel activity. He found Chad in central Africa to be a particularly interesting place to visit. The country is known as the crossroads of Africa as it occupies a strategic position in the centre of the continent. The north of the country is largely desert whereas the south is a more fertile savanna zone. Dr Wells travelled across the Sahara in temperatures of 136 degrees Fahrenheit to collect samples from the remote tribes. Wherever possible blood samples are taken from indigenous peoples because more DNA can be extracted from blood, and it is not known if an opportunity will ever arise again to visit these remote places. For the public participation programme a simple cheek swab is required. For the most part the local population are thrilled to participate in the research and are fascinated to learn more about their history through their DNA. There have however been problems in countries which were once under colonial rule, especially where land rights are involved, and the project is working closely with Native Americans and Aborigines to increase their participation.

At the end of the lecture there was a very lively question and answer session, and it was clear from the questions that the subject had inspired the public interest. Dr Wells was available after the talk to sign copies of his book Deep Ancestry: Inside the Genographic Project. He also told us that he has a new book due out in June 2010 entitled Pandora's Seed: the Unforeseen Cost of Civilisation which will focus on society and culture rather than genetics. Dr Wells is now starting to work on The Genographic Source Book, a huge compendium of all the data generated from the project, which is scheduled to be published in 2011.Further information can be found on the Genographic Project website. Public participation kits can be purchased in the UK from the National Geographic online store for £68.94 plus £4.95 for postage and packing. Kits are also on sale at the National Geographic Shop at 83-97 Regent Street, London, W1B 4E1, but are much more expensive at £99 (the same price in sterling as the retail price in dollars in the US!). Not surprisingly, therefore, very few of the people attending the lecture actually bought a kit on the day. If you are interested in purchasing a kit I would therefore recommend ordering direct from the National Geographic online store. Outside the UK, kits can be ordered direct from the Genographic Project website in America. The shipping costs are however very expensive for anyone not living in the US or Canada. For many people it will be more economical to test first through a surname or geographical project at Family Tree DNA and then transfer their results to the Genographic Project. To do so visit your FTDNA personal page, click on the Genographic Project link under Tools and follow the instructions. You will be asked to agree to the Project's consent terms, and there is a nominal fee of US $15 per test. Proceeds from this fee will be directed to the Legacy Project. For those people who test first with the Genographic Project I would recommend transferring your results to the Family Tree DNA database, where you can join the relevant surname, geographical and haplogroup projects, and order upgrades and further tests as required.
The Genographic Project will test either your mitochondrial DNA, which is passed down each generation from mother to child and reveals your direct maternal ancestry; or your Y chromosome (males only), which is passed down from father to son and reveals your direct paternal ancestry. I've already had my own mitochondrial DNA tested through Family Tree DNA, and have added my results to the Genographic Project database. I shall follow the progress of the project with interest and shall look forward to reading the research papers as they are published.