Showing posts with label Autosomal DNA. Show all posts
Showing posts with label Autosomal DNA. Show all posts

Monday, 22 January 2018

Small segments and pile-ups - a visualisation

We've recently been discussing the problem of pile-ups in the All Genetic Genealogy group on Facebook. A pile-up is a term used in genetic genealogy to describe multiple shared autosomal DNA segments that are stacked up on top of each other on the same part of the genome. The presence of a pile-up should be considered as a warning sign. For any shared segment to have genealogical significance we would expect it to be shared only with descendants of the common ancestral couple. If we share a segment with hundreds or thousands of people it is extremely unlikely that we will share that section of DNA by virtue of a recent genealogical relationship within the last ten generations or so, and it is much more likely to be indicative of a false match or a more distant relationship.

Pile-ups can occur for a number of different reasons:
  • Lack of phasing. Phasing is the process of sorting the DNA letters (the As, Cs, Ts and Gs) onto the paternal and maternal chromosomes. AncestryDNA and MyHeritage now used phased matching which means that they phase our genotypes before trying to identify shared sections of DNA. 23andMe and Family Tree DNA use a process of half-identical matching. Our DNA is not phased but instead the algorithms zigzag backwards and forwards across two columns of unsorted DNA letters looking for consecutive runs of matching SNPs. Half-identical matching works well at identifying large shared segments of DNA but is less successful on smaller segments, and particularly segments under about 10 centiMorgans (cMs) in size. if a match does not survive phasing it is a false match.
  • SNP-poor regions. The autosomal DNA tests used for genetic genealogy provide information on between 630,000 and 700,000 genetic markers known as SNPs (single nucleotide polymorphisms) which are scattered across the genome. These SNPs are only a tiny fraction of the three billion letters which make up the human genome, but the SNPs are specially selected for being the most informative about variations within and between populations. When trying to identify shared regions of the genome the companies are looking for long runs of consecutive SNPs that are the same (identical by state or IBS) in two individuals. Segments which pass the companies' matching thresholds are declared to be identical by descent (IBD) and are possibly indicative of shared ancestry in a genealogical timeframe. Some companies will also apply additional algorithms to filter out known problematic regions which are unlikely to be IBD. However, because not all of our SNPs are being tested, the length of a segment can be falsely inflated. One hypothesis is that lots of small segments can become conflated into longer segments. (1) This problem is particularly likely to occur in sections of the genome which have poor coverage on the chips. (2) 
  • Excess IBD. This is a term used to describe sections of the genome which are known to be widely shared in humans or in certain populations. Such regions often offer some type of evolutionary advantage. For an overview of known excess IBD regions see the section on excess IBD sharing in the ISOGG Wiki article on IBD. In addition to looking at the size of a shared segment, some IBD detection algorithms will, therefore, also take into account the frequency of the segment. (3) The more people who share a segment, the older it is likely to be. AncestryDNA apply their proprietary Timber algorithm to phased segments and they downweight the cM count for segments that are widely shared in their database. (4)
Each individual has their own personal pile-ups. It can be instructive to map out your pile-ups so that you are aware of your own danger zones. I've previously used Don Worth's ADSA (autosomal DNA segment analyser) tool which is available from DNAGedcom to look at my pile-ups. I've also use the matching segment search at GEDmatch (this tool is available to Tier 1 subscribers). (5)  These tools are very useful for identifying problems in specific regions but it's difficult to get a good idea of the bigger picture.

Following on from our discussion in the All Genetic Genealogy Facebook group, Dan Edwards has been working on an exciting tool to provide a new way of visualising pile-ups. It's possible that the tool will eventually be made available on the web but for the moment it is a bespoke service. Dan has been experimenting on some of my data. He has produced for me some charts showing the distribution of shared segments across my 22 autosomes and on the X-chromosome. Dan has kindly given me permission to share my charts which are reproduced below.

The charts are based on my Family Finder chromosome browser data from Family Tree DNA. FTDNA updated their match thresholds in May 2016, but they are still the only company that continue to include small segments under 6 cMs when inferring a relationship. It is generally accepted by genetic genealogists that the use of such small segments is problematical. (6)

The problem with small segments can be clearly seen in the charts below. Rather than being distributed evenly across my genome, the smaller shared segments form huge spires and skyscrapers. As the segment size increases the pile-ups are greatly reduced, but there are still some parts of my genome which have some quite sizeable pile-ups on segments over 10 cMs in size. Chromosomes 9, 14, 18 and 19, in particular, seem to have a few problem areas which it is probably best for me to avoid. As more matches come in, these spires and skyscrapers can be expected to grow even more. Remember too that FTDNA only reports "matches" on small segments if the match thresholds have already been met. If matches were reported on all matches in the database down to 1 cM it's likely that the spires would be even more pronounced.

If Dan is able to develop his tool further and make it more widely available it will be interesting to see how other people's pile-ups compare with mine. I hope that we might also be able to identify a reason for some of the pile-ups. In the meantime I hope you enjoy looking at my pictures.























Footnotes

(1) See: Chiang CWK, Ralph P, Novembre J (2016). Conflation of short identity-by-descent segments bias their inferred length distribution. G3 Genes Genomes Genetics 6: 1287.

(2) For a useful overview of SNP coverage on the chips used by AncestryDNA and 23andMe see Rebekah Canada's series of articles on the subject of exploring microarray chips.

(3) For a good overview of the methodology of IBD detection see Browning and Browning (2012):  Identity by descent between distant relatives: detection and applications (Annual Review of Genetics 2012; 46: 617-33). The authors state: "The key idea behind IBD segment detection is haplotype frequency. If the frequency of a shared haplotype is very small, the haplotype is unlikely to be observed twice in independently sampled individuals, so one can infer the presence of an IBD segment. This criterion can be applied in several ways. The first is length of sharing, which is a proxy for frequency. If two densely genotyped haplotypes are identical at all or most (allowing for some genotyping error) assayed alleles over a very large segment of a chromosome, then the haplotypes are likely to be identical by descent across the whole segment. The second is direct use of haplotype frequency: Shared haplotypes with estimated frequency below some threshold are determined to be identical by descent. The third makes use of a population genetics model to infer probability of IBD. Given the frequency of the shared haplotype and a probability model for the IBD process along the chromosome, one can estimate the probability that the individuals are identical by descent at any position on the segment."

(4) For a good explanation of how the AncestryDNA algorithm works read the blog post by Julie Granka on Filtering DNA matches at AncestryDNA with Timber. Take a look in particular at the figure in that blog post. Although the majority of phased segments filtered out by Timber are smaller segments under 15 cMs, note that it also downweights some larger segments up to 50 cMs in size.

(5) Peter Alefounder has developed a tool known as the Geneal Segment Stacker but I've not yet had time to play around with it. There are further details in this thread in the ISOGG Facebook group.

(6) For an excellent summary on the current state of our knowledge on the subject of small segments see the blog post A small segment round up by Blaine Bettinger.

Further reading

Saturday, 29 July 2017

Comparing match tallies for family members with Family Tree DNA's Family Finder test

I've taken a look at the total number of matches for all my family members who have taken a Family Finder test at Family Tree DNA. I've also done a comparison with the data I extracted on 26th May 2016 just before Family Tree DNA updated their matching algorithms. The results are shown in the table below. Note that I have excluded immediate family members from the totals.

Relation Number of matches
 28 July 2017
Number of matches
 26 May 2016
% increase
Debbie 1217 592 51%
Debbie's dad 1344 643 52%
Debbie's mum 1038 495 52%
Debbie's husband 905 443 51%
Debbie's eldest son 1138 542 52%

If my matches are representative of the wider Family Finder database then there has been over a 50% increase in the size of the database in the last 14 months.

I've also looked at the number of matches I share with my parents and taken stock of the number of matches which don't match either parent.

I share 501 matches with my dad. Of these, 320 were assigned to the paternal side with FTDNA's Family Matching tool. The remaining 181 matches were in common with my dad but did not meet the threshold for Family Matching.

I share 402 matches with my mum. Of these, 276 were assigned to the maternal side with the Family Matching tool. The remaining 126 matches did not meet the threshold for Family Matching.

I therefore have a total of 903 matches (74%) which match my mum or my dad. However, this means that 314 of my 1217 matches (26%) do not appear in the match lists of either of my parents.

All the matches that don't match my parents have a longest segment under 15 cMs. This is the breakdown.

Longest block  Number
10-14 cMs 28
7-9 cMs 286

The last time I did a comparison of parent and child matches I found that 23% of my matches did not match either of my parents.

These matches are either false positives or false negatives but without further investigation it is not possible to tell.

Have you tested both of your parents at Family Tree DNA? What are your statistics?

Related blog posts

Tuesday, 24 May 2016

New match thresholds for Family Tree DNA's Family Finder test

As a "trusted blogger" I have been given advanced notice by Family Tree DNA of forthcoming changes to the match thresholds for the Family Finder autosomal DNA test. The changes are to be rolled out very soon once the final quality control checks have been run. An e-mail will be sent out to project administrators in due course. Here are the details I received from Family Tree DNA:
For several years the genetic genealogy community has asked for adjustments to the matching thresholds in the Family Finder autosomal test. After months of research and testing, we will shortly be implementing some exciting changes.

The current matching thresholds – the minimum amount of shared DNA required for two people to show as a match are:

● Minimum longest block of at least 7.69 cM for 99% of testers, 5.5 cM for the other one percent

● Minimum 20 total shared centiMorgans 
Some people believed those thresholds to be too restrictive, and through the years requested changes that would loosen those restrictions.

The following changes will be made to the matching programme.

● No minimum shared centiMorgans, but if the cM total is less than 20, at least one segment must be 9 cM or longer.

● If the longest block of shared DNA is greater than 9 cM, the match will show regardless of total shared cM or the number of matching segments.

The entire existing database will be rerun using the new matching criteria, and all new matches will be calculated with the new thresholds.

Most people will see only minor changes in their matches, mostly in the speculative range. They may lose some matches but gain others.
This is very welcome news. This was a change that many of us had asked for and it's good to know that Family Tree DNA have listened to us.

When setting a cut-off limit it is always difficult to get the balance right between false positive and false negative matches but the previous 20 cM threshold was problematic because all segments right down to 1 cM were included in the total. Family Tree DNA do not currently phase their data before assigning matches (sort the alleles into the maternal and paternal chromosomes) and we know that the vast majority of unphased small segments, particularly under 7 cMs, are false positives.(1) Some people were therefore declared as matches when most of the segments they shared were small pseudosegments, and they were unlikely to share a recent common ancestor. In contrast, some legitimate cousin matches were not showing up because they fell just below the threshold. Under the old system two cousins could potentially share a 15 cM segment but not have enough of the small pseudosegments to make up the 20 cM quota. Anecdotally it has been observed that the 20 cM threshold was a particular problem for people with African ancestry who tend to have fewer of these false coincidental matches on small segments.

Some people were advocating for Family Tree DNA to set the threshold at 7 cMs, but the 9 cM threshold is a sensible compromise. There is still a high false positive rate for unphased 7-9 cM segments, so this will ensure that the reported matches are more likely to be real.

It should also be remembered that, in the vast majority of cases, if you match on a single segment under 10 cMs you will not share a common ancestor within the last ten generations. Even matches of 10 cMs can be very distant.(2) One study found that fewer than 35% of IBD (identical by descent) matches of 10 cMs fall within the last ten generations, and over 30% of segments of this size date back over 20 generations.(3)

I've also noticed in my own data that a lot of the segments in the 7 to 9 cM range seem to fall into large triangulated groups. If these segments are real then this is an indication that they are in what are known as pile up regions. These are regions of the genome where lots of people match because they share the same ethnicity or for some other reason rather than because they share a single recent common ancestor.

Indeed, because of the difficulties in working with unphased segments under 10 cMs many genetic genealogists recommend focusing only on matches who share 10 cMs or more.

I hope to do a comparison of my before and after matches at Family Tree DNA and will be interested to see comparisons from other people, but this is a very welcome and positive change. Thank you Family Tree DNA!

Update
It was not clear from the original announcement but it has now been confirmed that all matches with a total cM count of 20 cMs with a longest segment of 7.69 cMs or more in size will still be reported. Blaine Bettinger has provided a very useful decision tree to clarify the situation in his blog post Family Tree DNA updates matching thresholds. It therefore seems unlikely that many people will lose matches. Note that FTDNA does include all small segments right down to 1 cMs in their match thresholds. Most of these smaller segments, and especially those under 5 cMs are just noise and are best ignored unless you are able to do phasing and very careful chromosome mapping by testing a large number of close family members and known cousins.

Update 25 May 2016
I have received further information about the forthcoming update in an e-mail sent out by Family Tree DNA to all their volunteer group administrators. Here is the relevant section:
We also slightly altered other proprietary portions of the matching algorithm that will, to a small degree, affect block sizes and total shared centiMorgans. These changes should have only marginal effects, if any, on relationships, generally in the distant to remote ranges. 
There’s a separate proprietary formula that is also applied to those with Ashkenazi heritage, but you can, of course, expect to have more new matches than those not of Ashkenazi heritage. 
Please keep in mind this change will not affect close matches, only distant and speculative ones. Some matches will fall off, others will be added. Most people will likely have a net gain of matches. 
Your myOrigins results may change slightly with the rerun, but we have not updated or changed myOrigins yet. We’ll let you know when that happens.

See also

Footnotes
1. See the statistics on false positive matches on the ISOGG Wiki page on identical by descent.
2. See the blog post by Steve Mount on Genetic genealogy and the single segmentOn Genetics, 19 February 2011.
3. See Figure 2 in the paper by Doug Speed and David Balding on Relatedness in the post-genomic era: is is still useful? Nature Reviews Genetics 2015 6: 33-44. 

Tuesday, 17 May 2016

AncestryDNA are to use a new chip for their autosomal DNA test

AncestryDNA have announced that with effect from this week they will be using a new chip for their autosomal DNA test. The announcement was made in a blog post by the Ancestry team Customer testing begins on new AncestryDNA chip published on 12th May.

Some genetic genealogists in the US were invited to attend a conference call with the AncestryDNA team where they were given the chance to ask questions about the changes. For further details read the following two articles:
I will update this list if any further articles are published.

Friday, 6 May 2016

AncestryDNA's updated matching algorithms - a before and after analysis

AncestryDNA rolled out their long-awaited new matching algorithms on Tuesday this week. This message will now greet you when you log into your AncestryDNA account.
Ancestry have provided a number of resources to describe the changes, all of which merit a close reading:
AncestryDNA have been able to make these improvements because they have such a massive autosomal DNA database. They have now tested nearly two million people. Their scientists have been able to exploit the power of this large database to provide new insights into relatedness and to improve the detection of genealogically relevant IBD segments.

The biggest change is an improvement in the phasing process, Phasing is the process of sorting out the DNA letters   the As, Cs, Ts and Gs  – and placing them on the maternal and paternal chromosomes. Phasing is important for ruling out false positive and false negative matches. AncestryDNA are now using a reference panel of more than 300,000 genotypes for their phasing. Previously they were using a "window" system for IBD detection which broke the large segments into too many small pieces. Now they are using a SNP-based system which provides more realistic results with fewer segments. Phasing can be done with reference panels with a high degree of accuracy  the error rate of Ancestry's Underdog phasing engine is less than 1%. The accuracy will increase as the reference panel grows in size.

The matching threshold has also been changed. Two people must now share a minimum of 6 cMs whereas the old threshold was 5 cMs. AncestryDNA have produced a revised table of confidence scores based on a new understanding of the amount of DNA shared between different relations.


Contrast the above scores with the old version of the chart below which, to my mind, was always overly optimistic, especially about the matches on segments under 20 cMs, the vast majority of which are actually shared with very distant cousins. (For more on this subject watch Dr Doug Speed's lecture Who's your cousin? Using DNA to determine relatedness which he presented at  Who Do You Think You Are? Live this year.)


Comparing matches before and after
I thought it would be an interesting exercise to compare my matches before and after the update. Unlike Family Tree DNA and 23andMe, AncestryDNA do not provide a facility for customers to download their match list. Fortunately Rob Warthen from DNAGedcom has provided a tool known as the DNAGedcom Client, which allows us to download all our data from Ancestry, including details of the shared cM count and the number of shared segments. I downloaded my list of matches on 19th April. I ran the DNAGedcom Client again on 4th May, and I've compared the two datasets to see how many matches I've gained and lost.

Here is a comparison of the number of matches I had before and after the update:

DateMatches4th cousinsDistant cousinsShaky leaf hintsCirclesNADs
4 May3423183405100
19 April3414283386100

There was a marginal increase in the number of matches, but a close analysis of these matches provides a different perspective. I actually lost 1169 (34%) of my matches. However, this is more than made up for by the fact that I have gained 1178 new matches.

This is a breakdown of the size of the segments I share with my matches before and after the update:

Date< 6 cMs6-6.99 cMs7-9.9 cMs10-10.9 cMs>15 cMsMatches
4 May015181456381683423
19 April1737704719214403414

I thought it would be interesting to do a further breakdown of my matches who were predicted to be fourth cousins or closer. Note that what AncestryDNA describe as a fourth cousin can in fact be anything from a fourth to a sixth cousin.

Relationship beforeRelationship aftercMs beforecMs afterSegments beforeSegments after
3rd cousin3rd cousin109.71117.19785
3rd cousin3rd cousin84.35898.316944
4th cousin4th cousin52.36161.044543
4th cousin4th cousin23.94730.443543
4th cousin4th cousin23.69529.488911
4th cousin4th cousin21.77627.269111
4th cousin4th cousin24.09825.291522
4th cousin4th cousin22.32524.425911
Distant cousin4th cousin11.36423.877132
4th cousin4th cousin20.06523.863411
4th cousin4th cousin20.70423.197522
4th cousin4th cousin17.15922.907622
Distant cousin4th cousin13.20422.896221
Distant cousin4th cousin13.65322.273432
4th cousin4th cousin18.58121.195611
4th cousin4th cousin22.60420.964411
4th cousin4th cousin20.89620.848321
4th cousin4th cousin18.03320.650611
4th cousinDistant cousin24.06813.054511
4th cousinDistant cousin19.14919.370221
4th cousinDistant cousin18.62918.279611
4th cousinDistant cousin18.47418.610711
4th cousinNone24.5381
4th cousinNone23.7311
4th cousinNone23.0991
4th cousinNone21.9781
4th cousinNone21.2851
4th cousinNone21.0931
4th cousinNone20.8781
4th cousinNone20.1721
4th cousinNone19.9371
4th cousinNone18.0711

As can be seen, for the matches that have been retained there has been a marginal increase in the cM count. Four matches have been downgraded from fourth cousins to distant cousins. Ten of my previous fourth cousins (35%) have disappeared from my match list completely. It may be that these matches were filtered out because of the improved phasing. Another possibility is that these segments were in SNP-poor regions. Ancestry explain in their white paper that matches in these regions are unreliable. To counteract this problem they "discount these matches by reducing their total length (in cM)". These matches are no great loss. All these fourth cousins were in America and it was impossible to find any sort of genealogical relationship despite the fact that some of these matches had huge and very detailed trees. I'd rather suspected that these matches must be very distant, if they were legitimate at all. I already have far more matches than I know what to do with and I can still only find the genealogical connection with two of my matches at AncestryDNA. I would much rather have fewer and more accurate matches.

Conclusion
It's important to remember that we are all pioneers in this field, the tests are in their infancy and we still have much to learn.

At Family Tree DNA, 23andMe and GedMatch we are used to working with unphased data which can produce false positive matches, particularly on smaller segments under 15 cMs. Ancestry are the only company who filter out the high-frequency matches which are not of genealogical relevance, though 23andMe do screen out some matches in known problem areas. IBD segments with high rates of matching are likely to be less useful for detecting relationships in a recent genealogical timeframe.

Without phasing and without frequency filters it is much easier for people to find false coincidental matches, but we all need to be very careful about jumping to conclusions, especially with more distant relationships, where it is so much more difficult to detect recent IBD with the currently available tests.

This is the second time that AncestryDNA have updated their algorithms. Family Tree DNA have already changed their algorithms once, which resulted in some lost matches. We should all expect to see further changes to the companies' matching algorithms in the future as they strive to improve the technology and produce more accurate results.

Further reading
Thumbs up; AncestryDNA improves genetic matching technology - a review by Diahan Southard, 9 May 2016.

Acknowledgements
Thanks to Don Worth in the ISOGG Facebook group for sharing his Excel formula for calculating the number of lost matches.

© 2016 Debbie Kennett