Showing posts with label Tajik. Show all posts
Showing posts with label Tajik. Show all posts

Saturday, July 13, 2013

A Hidden Gem in Central Asia: Previously Unknown Y-DNA R1b Haplotype [Original Work]

1. Introduction

Central Asian Y-DNA diversity has been an area of constant intrigue in the genetics community. Wells et
al.'s The Eurasian Heartland: A continental perspective on Y-chromosome diversity paved the way, with several others following in their regard. Members of the same team (including Dr. Wells) produced another paper - A Genetic Landscape Reshaped by Recent Events: Y-Chromosomal Insights into Central Asia - on the same topic in the following year, this time headed by Dr. Tatania Zerjal. I noted a greater emphasis on East-Central Asian populations as well as a mentioning of Y-STR analysis in the study itself. However, none of this data was supplied, with only Y-SNP information included (shown sporadically in this entry). The age of this paper is apparent through the nomenclature used (see Method section).

Several months ago, I made a request to obtain the Y-STR data from this study to one of the co-authors, Dr. Tyler-Smith, who kindly replied with the results of all sampled populations (Data Sink > Zerjal et al. Raw Data).

In this blog entry, the Y-STR data is showcased with a special emphasis on the Y-DNA R1b-M269 which was discovered.


2. Method
Y-SNP Phylogeny in original paper (Zerjal et al.) [1]


The maximum number of compatible Y-STR's were utilised for processing in Urasin's YPredictor for easier haplogroup identification (14 of a possible 16, DYS434 and 435 were excluded). All data was run through YPredictor. Only samples with ≥70% probability were included in the final results (Data Sink > Processed Data). As discussed below, relevant findings are compared with the basic Y-SNP haplogroups shown in the original study (on right).

One point which needs to be addressed immediately is the high frequency of "_DE-M1" and "P-M45". It appears that the STR selection has led to a phantom result, rendering many of the samples useless. For instance, the original study shows the Kazakhs belong overwhelmingly to C3c-M48, [1] although the probable results shown here are mostly "_DE-M1".  The exclusion of DYS434 and 435 from my level of processing likely contributed to this; if one assigns equal weight to the statistical strength of a prediction, removal of two STR's from a panel numbering 16, accuracy is reduced by 12.5%. Additionally, some conversion error seems to have applied with DYS437 (i.e. a value <12 is unusual). Therefore, "_DE-M1" and "P-M45" results were dismissed on account of the mismatch between predicted and likely confirmed haplogroups probably due to a compatibility issue between the study's STR panel and YPredictor..


3. Results

As the majority of samples were removed owing to the caveat described above, this entry will take a qualitative rather than quantitative approach to analysis on the general picture formed. Much of the remaining results are congruent with findings in other papers. Populations around the Caucasus are signified by plenty of R1b-M269, J2a-M410 and G2a-P15. Tajiks and the Kyrgyz were predominantly R1a1a-M17. Mongolians and other East-Central Asian ethnic groups yielded the most O3-M122 and "NO-M14" (likely to be Y-DNA N or O suffering from the STR restrictions described in the Method section).


Y-SNP distribution in Central Asia (Zerjal et al.) [1]


3i. The R1b Signal 

R1b-M269 was found across Central Asia and not only in the Caucasus (Armenians, Azeris, Georgians, Ossetians). It was mostly detected among the Turkmen (trk1, trk2, trk4, trk6, trk7, trk22, T29, T32) with a single sample among the Uzbek (uz-s110). [1]

Analysis of the haplotypes (including DYS434 and DYS435) revealed the nine Central Asian R1b samples belonged to a secure haplotype (Data Sink > R1b Results). trk6 diverged greatest, albeit with two 1-step mutations on DYS393 and DYS434. The rest match this haplotype exactly or have single 1-step mutations. [1] When this Central Asian R1b haplotype is compared with the other Caucasian samples, a mixed picture emerges, with the poorest being an Armenian (arm47) at 8/16, whereas the best are another Armenian (arm12) and Azeri (az48), both at 15/16. [1]
One interesting point is the Kurds sampled in this study (some of whom also belong to R1b-M269) are actually the displaced population positioned on the Iranian-Turkmenistani border. All of whom match the Central Asian R1b haplotype with a similar value (12-13/16). This definitively rules out the Kurds as a source for the haplotype, particularly as better matches can be found further to the west. It should be noted the Kurds themselves formed their own R1b haplotype (defined here by DYS389II=27, DYS391=10). [1]

In summary, the data reveals that the Turkmen are particularly abundant in R1b-M269 and all belong to the same haplotype as one of the Uzbek samples. This haplotype matched some Caucasians very well, but others not so well. The Kurds living in Turkmenistan belonged to their own haplotype.


3ii. Is This Actually R1b-M269?

Attention must first be shown to the original paper again; any potential R1b-M269 here will be present as P(xR1a)-92R7 (shown in the paper as "Haplogroup 1"). [1] Evidently, this makes up approximately half of the Turkmen lines and a quarter of Uzbek ones. Other haplogroups (such as other forms of R1b, R2a-M124, various Q subclades) presumably make up the rest of "Haplogroup 1" shown.

The next step is to verify whether or not this Central Asian R1b haplotype matches other R1b haplotypes online. As Y-DNA R1b-M269 is fortunately well-represented in the world of genetic genealogy, searching for the haplotype's matches on ySearch is a reasonable enterprise. DYS437 had to be excluded here due to a conversion issue, leaving the haplotype at 15 STR's. A genetic distance (GD) of 3 was allowed on these 15 markers. Results are shown on the right.

ySearch results for Central Asian R1b haplotype
With some confidence, the search has demonstrated that the Central Asian R1b haplotype does indeed belong to R1b-M269, as all the seven matches shown (one of whom is Armenian) belong to it.

Expanding the line of inquiry one further step came through comparing this haplotype with Iranian haplotypes [2] which were readily available. Due to differences in STR panels (an overlap of only 11) this proved to be inconclusive, aside from the observation that DYS389i+ii was completely different between the Central Asian modal (10-26) and the Iranian values. At this point I suspect that, much like DYS437, there is a conversion issue with DYS389 also.

Finally, a comparison was made with the R1b found in Afghanistan last year [3]. Interestingly, if DYS389i+ii and DYS437 are excluded, the two Uzbeks (samples 35 and 181) match the Central Asian R1b haplotype almost exactly based on the remaining 11 STR's. The one Tajik (sample 32) is less likely to be related due to two 1-step mutations on different STR's.


4. Conclusion

The inferences made from the data hang by a metaphorical thread due to the persistent STR issue; different labs have used different panels in the past decade, making it excruciatingly difficult to use materials from older papers. Fortunately, the presence of a specific strain of R1b-M269 in Central Asian (in Turkmen and Uzbeks) has successfully been demonstrated after select exclusions and no modifications to the data.

However, some larger questions remain. If STR limitations were not an issue, how would the Iranians from Haber et al. have compared? Would the Tajik from the other Haber et al. paper have belonged to the same haplotype in the end?

The origin of this Central Asian R1b haplotype will, I anticipate, also be a point discussed heavily among interested parties. At this point in time, I must stress that none of the evidence thus far points to anything in particular without ruling other theories out, although it leaves the door for interpretation wide open.

Having given this cautionary statement, the main thrust of this entry should be emphasised; R1b-M269 in Central Asia is a confirmed reality and here to stay. I will defer any subsequent analyses to the experts on Y-DNA R1b which grace several genetic genealogy boards for their take on the flavour of this haplotype.


5. Acknowledgement

I publicly extend my gratitude to Dr. Tyler-Smith for being so kind in sending me the raw STR's from this important paper for my research, as well as co-authoring the other two excellent studies I have cited here and in the past.


6. References

1. Zerjal T, Wells RS, Yuldasheva N, Ruzibakiev R, Tyler-Smith C. A genetic landscape reshaped by recent events: Y-chromosomal insights into central Asia. Am J Hum Genet. 2002 Sep;71(3):466-82. Epub 2002 Jul 17.

2. Haber M, Platt DE, Badro DA, Xue Y, El-Sibai M, Bonab MA. Influences of history, geography, and religion on genetic structure: the Maronites in Lebanon. Eur J Hum Genet. 2011 Mar;19(3):334-40. doi: 10.1038/ejhg.2010.177. Epub 2010 Dec 1.

3. Haber M, Platt DE, Ashrafian Bonab M, Youhanna SC, Soria-Hernanz DF, Martínez-Cruz B. Afghanistan's ethnic groups share a Y-chromosomal heritage structured by historical events. PLoS One. 2012;7(3):e34288. doi: 10.1371/journal.pone.0034288. Epub 2012 Mar 28.

Saturday, December 22, 2012

Yaghnobi Tajiks: Preliminary Results May Reveal Iranian Plateau Affinity [Original Work]

Slipping under the radar of the genetic genealogy world is this paper by Elisabetta Cilli and her colleagues, which investigated the mitochondrial data of 62 individuals from Tajikistan's Yaghnobi population. [1]

The Yaghnobis are of interest given their geographical isolation and the East Iranic nature of their language. Living just northeast of the predominantly Persian (Dari) speaking capital, Dushanbe, Yaghnobi is a continuation of a fully agglutinative Soghdian dialect representing the sole survivor of this language following the Persianization of Central Asia in Medieval times [2]. Despite its' East Iranic vocabulary, Yaghnobi demonstrates several linguistic features (i.e. gender loss, past imperfective preservation from present stem of a verb) which separates it from those modern East Iranic languages immediately surrounding it. Furthering the uniqueness of the Yaghnobi language in this context is the unity it forms through these features with languages mostly spoken further west in the Iranian plateau (e.g. Persian, Gilaki, Kurdish dialects). [2]

Although the results are preliminary and lack any empirical data, Cilli et al. have discovered some interesting connections between the Yaghnobi and relevant populations. In summary, they found the following:

MDS Plot of Results
  • 42 individuals used for the preliminary work belonged to only 19 distinct mtDNA haplotypes. Of these, 11 were distinct among the Yaghnobi.
  • The Yaghnobi have less mtDNA genetic diversity than other Central Asian populations (0.930) and this is attributed to their geographical isolation and recent history of displacement by the U.S.S.R. in the 1970's for agricultural purposes, where a small group (300) returned and repopulated their original homelands.
  • Intriguingly, the Yaghnobi shared all of the mutual haplotypes (8/19) with populations from Iran (e.g. Gilakis, Mazandaranis and Iranians from Tehran and Esfahan) instead of other Central Asian groups, including their Tajik compatriots.
  • The Yaghnobi shared most of these mutual haplotypes with Gilakis, Kurmanji Kurds and Avars from the Caucasus (4 each).
  • However, owing to their predominantly distinct mtDNA character, the Yaghnobi are clear outliers from the general zone occupied by the reference groups. 

My critique and interpretation of these results are as follows:

  • At least two instances of genetic drift occurring (founder effect via geographic isolation, bottleneck due to Soviet relocation) is likely responsible for the decreased mtDNA diversity. Thus, it is clearly simply a reflection of their environment.
  • As a result of the Soviet relocation, it may be useful to determine whether results from the displaced parent population match what has been stated here. This is quite possible given the relocations occurred just over one generation ago (~40 years).
  • It is difficult to criticise the decision to test 62 individuals and the utilisation of 42 haplotypes, given the Yaghnobi population in their homeland between 2007-9 only numbered approximately 500. Approximately 8% of the entire Yaghnobi population was therefore analysed here, which is a generous frequency given the amount of attention the region has received.
  • The MDS plot would have benefited from the inclusion of populations in Europe, Southwest Asia and South Asia to comprehensively flesh out the position of Yaghnobis in Eurasia.
  • Accepting that this is a preliminary investigation, it would still have been pleasing to see some raw data published. Aside from confirming that some/one Yaghnobi matched the Cambridge Reference Sequence (CRS, thus Haplogroup H2a2a which happened to be found in all the populations tested), there is no indication as to what the other mutations looked like. Or, for that matter, what mtDNA haplogroups were even present!


Correlation with Y-Chromosomal Data?

The Yaghnobi have been studied at least one other time through their inclusion in Dr. Spencer Wells et al.'s seminal piece The Eurasian heartland: a continental perspective on Y-chromosome diversity. The breakdown of their Y-Chromosomal SNP data (n=31) is as follows: [3]

3% C-M130(xC3a3-M48)
32% J2-M172
Y-SNP clustering reveals Yaghnobis sit near SE Europe and the Near-East
3% K-M9(xO-M175, O3-M122, O1a-M119, O2a1-M95, N1c1-M46) (possibly parahaplogroup such as K*-M9)
10% L-M20
3% P-M45 (xQ1a1-M120, Q1a3a1-M3, R2a-M124)
32% R1-M173 (likely R1b1a1-M73 or R1b1a2-M269)
16% R1a1a-M17(xR1a-M87, private marker)

Despite the double genetic drift undoubtedly affecting the frequencies, it is worth pointing out that the Yaghnobi presented with a broadly similar Y-DNA spectrum as Iran, where J2-M172, L-M20, R1-M173 and R1a1a-M17 (including subclades) comprise approximately 53% of the national average (refer to Grugni et al. analysis). 

This comparison should be taken with a grain of salt given the Iranian national average also comprises non-Iranic-speaking ethnic groups, the Wells Yaghnobi data does not present with thorough downstream Y-SNP evidence, the sample size is contentious and at least two contributors of a founder effect exist. However, that the Yaghnobi appear rich in J2, L and R is certainly reminiscent of Iranic-speaking populations in the region.


Conclusions

The Yaghnobi are an exceedingly interesting population whose overall parental markers seem to support a connection with populations further west than one would anticipate.

Despite the misgivings of all the data concerning them to date, the mtDNA similarity does corroborate specific linguistic features between the Yaghnobi language with those in the Iranian plateau, such as Kurdish or Persian.

If the data holds up in future investigations, it certainly calls to question whether the proposed model of linguistic inheritance exclusively down the parental line (as represented by Y-DNA data) is entirely correct given this connection.

How the Yaghnobi came to display the markers within them whilst speaking an East Iranic dialect with traits akin to those found in West Iranic languages is an intriguing question. One possible scenario is that the Yaghnobi are partly descended from ancient Iranians from the Iranian plateau during the Achaemanid era. This would also account for the linguistic commonalities noted in current literature.

Time (with the assistance of more mtDNA, Y-DNA and auDNA) will help us understand what happened in Central Asia during the formative period that was the Indo-Iranian migrations.



Reference

1. Cilli E, Delaini P, Costazza B, Giacomello L, Panaino A, Gruppioni G. Ethno-anthropological and genetic study of the Yaghnobis;an isolated community in Central Asia. A preliminary study. J Anthropol Sci. 2011;89:189-94.

2. Windfuhr, G. The Iranian Languages. 1st ed. Routledge Language Family Series. 2009.

3. Wells RS, Yuldasheva N, Ruzibakiev R, Underhill PA, Evseeva I, Blue-Smith J. The Eurasian heartland: a continental perspective on Y-chromosome diversity. Proc Natl Acad Sci U S A. 28;98:10244-9. 2001.

Thursday, March 29, 2012

Showcasing of Y-DNA Variation Among Afghan Ethnic Groups [Review]

This very recent paper on Afghan Y-Chromosomes was released by M Haber et al. and provides us with an insight into the paternally-determined genetic structure of several Afghan populations.

Afghanistan's Ethnic Groups Share a Y-Chromosomal Heritage Structured by Historical Events
Haber M, Platt DE, Ashrafian Bonab M, Youhanna SC, Soria-Hernanz DF, et al. (2012) Afghanistan's Ethnic Groups Share a Y-Chromosomal Heritage Structured by Historical Events. PLoS ONE 7(3): e34288. doi:10.1371/journal.pone.0034288

"Afghanistan has held a strategic position throughout history. It has been inhabited since the Paleolithic and later became a crossroad for expanding civilizations and empires. Afghanistan's location, history, and diverse ethnic groups present a unique opportunity to explore how nations and ethnic groups emerged, and how major cultural evolutions and technological developments in human history have influenced modern population structures. In this study we have analyzed, for the first time, the four major ethnic groups in present-day Afghanistan: Hazara, Pashtun, Tajik, and Uzbek, using 52 binary markers and 19 short tandem repeats on the non-recombinant segment of the Y-chromosome. A total of 204 Afghan samples were investigated along with more than 8,500 samples from surrounding populations important to Afghanistan's history through migrations and conquests, including Iranians, Greeks, Indians, Middle Easterners, East Europeans, and East Asians. Our results suggest that all current Afghans largely share a heritage derived from a common unstructured ancestral population that could have emerged during the Neolithic revolution and the formation of the first farming communities. Our results also indicate that inter-Afghan differentiation started during the Bronze Age, probably driven by the formation of the first civilizations in the region. Later migrations and invasions into the region have been assimilated differentially among the ethnic groups, increasing inter-population genetic differences, and giving the Afghans a unique genetic diversity in Central Asia."

[PDF] [Supplementary Data]

Tabulated Y-DNA Haplogroup frequencies of the 204 individuals sampled distinguished by ethno-linguistic affiliation (ISOGG 2011 Nomenclature utilised) can be found in the Data Sink.



Results (populations sample count ~50 only)


- Haplogroup B-M60, a marker that would normally be expected among African populations, makes a surprising presence in the Afghan Hazara. Superficial STR analysis (17/19 haplotype match between all) suggests a recent common paternal ancestor, although the timeframe and ultimate origin of this common ancestor is another question.

- Haplogroup C3-M217 has invariably been associated with the expansion of Altaic-/Mongolic steppe populations since medieval times. The greater frequency (33.9%) in the Hazara relative to the Tajiks and Pashtuns appears to support this, as well as the commonly-held belief they partially descend from Mongolian tribes.

- The Hazara E1b1b1c1-M34 also stems from a common ancestor (all three share the exact 19 STR haplotype).

- The single man belonging to Haplogroup G1-M285 is of Tajik descent. It is possible this man's paternal line arrived with eastward migrating Persians following the Sassanid collapse in 651 A.D.

- As shown in previous studies, the Pashtun Haplogroup G men are again G2c-M377 (entirely this time, in contrast with Lacau et al.)

- Paragroups H*-M69, J2a*-M410, Q*-M242 and R*-M207 all indicate that Afghanistan played an important role in the demic development of their downstream subclades, or was at the very least a geographic nexus. It is worth noting that the Hazara Q* men belong to a different haplotype to their Pashtun and Tajik compatriots, again indicating genetic drift has taken place since the formation of the Hazara ethnic group (or, instead, paternal consistency through the presumed Mongolic layer that eventually formed modern Hazaras).

- In previous studies (Sengupta et al., Lacau et al.), several haplotypes without backbone SNP testing were found to belong to Haplogroup I, which is frequently considered a lineage specific to Europe. For the first time we have evidence of an I clade (I2b1-M223) in South-Central Asia, specifically among the Hazara and Tajik. The following is a recent exchange with Professor Ken Nordtvedt regarding the I2b1-M223 samples;

"The two Hazara seem related.  Both haplotypes look like M223+, with the Tajik one like Continental2 characteristic of central Europe.
The Hazara haplotype looks more like M223+ Roots.  But both have some problems with being considered close matches to European haplotypes.
...
I don’t think such tmrcas would be worth much.  I still don’t have a firm subclade of M223 to work with for either haplotype."

Due to the limited STR's it is not possible to cleanly place these I2b1 haplotypes into any of the existing clusters/subclades. However, Haplogroup I2b itself appears to be thousands of years old (Nordtvedt's I tree, final page). This opens up the possibility for an endogenous form of Haplogroup I existing in South-Central Asia.

- A single Tajik belonging to J1c3-P58 was postulated to potentially be of Arabian origin. As the (miniscule) Afghani Arabs did not yield any J1c3, other possibilities should be considered, such as contacts with the Iranian plateau over the past few millenia.

- The Tajiks were the only population to boast the presence of all major subclades within Haplogroup L (L1a-M27, L1b-M317, L1c-M357). In line with their greater frequency relative to the Tajiks and Hazaras, several Pashtun L1c-M357 samples share similar (exact-to-2-step mutation) matches, suggesting another example of genetic drift.

- Although the Laghman Pashtuns share a similar L1c-M357 haplotype (16-17/19 match), so does the sole Tajik L1c from the same location, providing us with genetic evidence of recent mutual origins between Pashtuns and Tajiks in certain parts of Afghanistan.

- The Tajik population is more paternally diverse than all others sampled. Explanations include a less endogamous cultural character or the more recent imposition of the "Tajik" identity, which arrived with the medieval Turks.

- R1b1a*-P297 (xM269) and R1b1a2*-M269 (xU106) both appear in Uzbek and Tajik populations. Both the R1b1a*-P297 haplotypes are identical and belong to a Tajik and Uzbek, again showing there is some recent paternal overlap between Central Asian ethnic groups. I discovered the haplotype does not generally correspond with any of the established clusters in the R1b1a1-M73 Project, although there is a 13/15 match with a Tajik from Cluster B1. Although the limited STR's are unfavourable, I am of the opinion the match is substantial and the R1b1a*-P297 reported in this study is in fact R1b1a1-M73 and belongs to Cluster B1, whose membership also consists of other Tajiks, Uzbeks and an Anatolian Turk.

- It is very interesting to note that all the locations showing R1b1a*-P297 (xM269) and R1b1a2*-M269 (xU106) (Badakhshan, Herat, Takhar and Mazar-e-Sharif) lie on a horizonal plane that runs across the north of Afghanistan, particularly as the Bactria-Margiana Archaeological Complex (BMAC) was situated here.


Criticisms of Paper

- Haplogroup R2a-M124 has been erroneously correlated with aboriginal Subcontinental populations when results from the R2 WTY Project indicate places like India are a "sink" rather than a "source" (most Indian R2a is R2a1-L295, which has a spotty distribution across the rest of Eurasia).

- Haplogroup L is, much like R2a, an understudied lineage, presumably due to its' paucity in Europe. The once-common assumption in the population genetics and genealogical world that the frequency of a given lineage in a region/population signifies its' antiquity there has been proven to be inherently false through STR and SNP analysis. Haplogroup L may enjoy greater frequencies in India according to the sources at their disposal, but the presence of different L subclades in Central and West Asia should have at least given the authors the initiative to investigate the lineage's deeper structure rather than relying on a population genetics tagline from at least 2006 (Sengupta et al.).

- Despite the recent boon in research on Haplogroup R1a1a-M17's structure by independent genetic genealogists and projects (such as the R1a1a and Subclades Y-DNA Project), Haber et al. failed to include any of the pivotal SNP's that have been discovered since Underhill et al. from 2009, thus preventing observers from making any meaningful conclusions from the current findings, particularly in the context of the Indo-European migrations (generally accepted from the Eurasian steppes).

- When divided into ethno-linguistic lines, this study showcases 3 Arabs, 13 Balochis, 59 Hazaras, 5 Nurestanis, 49 Pashtuns, 56 Tajiks, 1 Turkmen and 17 Uzbeks. The most immediate criticism is inadequate testing of the Arabs, Nurestanis, Balochis, Turkmens and Uzbeks in particular.


Evaluation

Despite several glaring flaws in methodology, Haber et al. has provided us with a much-needed insight into the deeper genetic structure of Afghanistan's Y-Chromosome diversity. There is clear evidence of genetic drift (particularly among the Pashtun Q*-M242/L1c-M357 or Hazara C3-M277), as well as evidence of recent line sharing between populations (The situation of L1c-M357 in Laghman).

However, Haber et al. has thrown out some very interesting surprises (T1-M70 among Tajiks only) as well as validating results from previous studies that had previously been questioned (I2b1-M223 and R1b1a2-M269 particularly). How did these lineages arrive in Central Asia? Is recent colonial admixture a possibility? For the time being, we will have to contend with this questions steadfastly.


Addenum I [30/03/2012]: Determination of R1b1a*-P297 furthered with regard to it potentially being R1b1a1-M73.
Addenum II [30/03/2012]: Insertion of Nordtvedt correspondence.