Annual Retreat 2026

September 4th, Friday @ Monadnock Room (2nd fl), Broad Institute (415 Main Street, Cambridge, MA 02142)

- Registration is required -

Reg QR 2026 BEGS AR

 

 


8:45AM - Breakfast 


9:00AM - Welcome/Introduction [Organizers]


9:10AM - 1. David Cheek: Naxerova Lab  |  Harvard Medical School

  • Title: Cell death is a pervasive driver of cancer risk
  • Summary: How carcinogens cause cancer without inducing mutations remains largely unclear. We propose that many carcinogens drive cancer risk simply by killing cells: cell death triggers compensatory proliferation, promoting clonal expansion of pre-existing cancer-causing mutations. A computational multi-stage model predicts that cell death has a carcinogenic effect that rivals or exceeds mutagenesis. This is consistent with human somatic mutation data across various tissues and carcinogens, including smoking, alcohol, and inflammation. In mice, alcohol kills hepatocytes and promotes clonal expansions in liver without increasing mutational burden. Chemotherapy promotes pre-cancerous clonal expansions in blood, with the death of stem – but not differentiated – cells explaining leukemia risk. Together, our results point to cell death as a pervasive driver of carcinogenesis, with implications for predicting cancer risk and identifying carcinogens.

9:30AM - 2. Irwin Jungreis: Kellis Lab |  MIT

  • Title: ORBL: Measuring conservation and constraint on “ORFness” of non-canonical ORFs
  • Summary: ORF Relative Branch Length (ORBL) measures cross-species evolutionary conservation and constraint on the “ORFness” of an ORF, without regard to conservation of the encoded amino acid sequence. It is intended to detect ORFs encoding functional but poorly conserved peptides, as well as ORFs whose translation is functional but which do not necessarily encode a functional peptide, such as regulatory uORFs. This is particularly relevant for non-canonical ORFs (ncORFs). ORBLv measures conservation of ORFness by calculating the relative branch length of the phylogenetic tree of species in an alignment that have an intact orthologous ORF. ORBLq measures evolutionary constraint on ORFness by calculating the quantile of its ORBLv score among the ORBLv scores of untranslated ORFs of the same biotype and similar length, which corrects for conservation due to chance or to constraint on an overlapping coding sequence. We find that among a set of 7264 non-canonical human ORFs having Ribo-seq evidence of translation, a substantial fraction have ORBLq evidence of evolutionary constraint on ORFness, particularly uORFs and uoORFs, despite almost none having PhyloCSF evidence of conservation at the amino acid level. ORBLq scores correlate with other evidence of translation and function, including HLA peptide detection and CRISPR essentiality.

9:50AM - 3. Daniel Tabin: Reich Lab  |  Harvard

  • Title: Analysis of a ~31,000 year old genome from the Tuyana site reveals spread of Aurignacian ancestry to far eastern Siberia
  • Summary: We generated genome-wide data from a man who was buried around ~31,000 years ago at the Tuyana site in far eastern Siberia. By co-analyzing Tuyana’s data alongside previously published data, we find that he shared distinctive genetic ancestry with an Aurignacian culture individual who lived ~39,000 years ago in Belgium in far western Europe, lending genetic support to the surprising archaeological claim that that the Tuyana individual was associated with the Aurignacian culture. This Aurignacian ancestry signature was also present in later Paleolithic Siberians including the Salkhit individual from Mongolia who lived ~35,000 years ago and for whom we report the first shotgun sequencing data, Yana individuals who lived ~32,000 years ago in northeast Siberia, and the Mal’ta individual who lived in ~24,000 years ago near Lake Baikal, highlighting how the spread of Aurignacian culture had an impact not just archaeologically but also genetically across large parts of northern Eurasia. However, Mal’ta, Yana, and Salkhit did not descend directly from Tuyana’s population, as Tuyana also had ancestry related to the Ust’-Ishim individual from western Siberia who lived ~45,000 years ago and who was previously viewed as deriving from a dead-end Initial Upper Paleolithic expansion into Eurasia. We also find evidence that the ancestry shared by Mal'ta and Balkan Hunter Gatherers comes from a shared ancestry of West Eurasian, rather than East Eurasian origin, and postulate that this population may have lived just West of the Caspian Sea

10:10AM - Break


10:30AM - Panel Discussion


11:40AM - 4.Kepler Mears: Baym Lab |  Harvard Medical School

  • Title:  RNA-guided nucleases enable the intracellular gene drive of insertion sequence in plasmids
  • Summary: Mobile genetic elements (MGEs) and the interactions between them are a major source of evolutionary innovation. Insertion sequences, the most simple MGEs encoding only the necessary genes for transposition and maintenance, are widespread in bacterial genomes, but are particularly common in plasmids. Plasmids, self-replicating extrachromosomal DNA elements, often exist in multiple copies imparting a stochastic barrier to the fixation of an insertion sequence by limiting the proportion of the plasmid population harboring the IS. In this work we demonstrate that to overcome this, the IS200/605 family of insertion sequences utilizes the programmable RNA guided nuclease TnpB as gene drive to spread the IS through the plasmid population. TnpB, the likely ancestor of Cas12, records the specific insertion site of the IS in its RNA guide to prevent loss of the IS during transposition. When introduced to a plasmid TnpB will be reprogrammed to target and cleave IS- plasmids, resulting in biased replication of IS+ plasmids. Furthermore the gene drive activity is critical for the IS to invade high copy plasmid populations, and provides an additional advantage in intercellular competition between conjugative plasmids. The advantage in plasmid occupancy further sheds light on the prevalence TnpB has become so prevalent across the tree of life despite being unable to engage in horizontal gene transfer. More generally, the unique pressures arising from movement between genetic contexts with different multiplicities shapes the evolution of strategies for MGE spread.

12:00PM - 5. Annabel Perry: Reich Lab  |  Harvard

  • Title:  plAIgue: An Artificial Intelligence Framework for Detecting Microbial Species in Ancient Human DNA
  • Summary: We introduce an artificial intelligence (AI) model, plAIgue, which accurately assigns a probability of whether an ancient human DNA sample came from a person who died while infected with a microbial species. Detecting pathogens which circulated in the blood of ancient humans has emerged as a new window to understand diseases which afflicted our species in bygone eras and to study pathogen-host co-evolution. Existing methods to detect pathogens in ancient skeletal remains, however, require a series of ad hoc filters and do not quantify detection uncertainty. We propose plAIgue to overcome these limitations. We evaluated plAIgue’s ability to detect the Black Death pathogen (Yersinia pestis) in DNA sequencing data from ~200 ancient humans. Depending on data curation stringency, plAIgue achieved 87-97% accuracy, 0.36-0.05 binary cross entropy loss, 1.0 sensitivity, and 0.95 specificity in our validation set. A state-of-the-art screening tool, the MEGAN ALignment Tool (MALT) followed by Heuristic Operations for Pathogen Screening (HOPS), achieves 0.27-1.0 sensitivity, and 0.47-1.0 specificity (depending on the HOPs step to which we compared) on the same dataset. These results suggest that plAIgue has the potential to perform as well or better than the gold standard method in a fraction of the time, with less human labor, and with quantification of detection uncertainty. Interpretability analyses show plAIgue uses variables and genomic regions indicative of Y. pestis infection when diagnosing true positives. We will generalize our method to other pathogens. Finally, analysis of more than 20,000 ancient individuals provides an opportunity to benchmark our model’s capabilities.

12:20PM - Lunch


1:20PM - 6. Dan Balick: Sunyaev Lab |  Harvard Medical School

  • Title:  Reanimating repeat length dynamics buried in the transposon graveyard: inference of repeat instability rates from the non-equilibrium response to transposon inactivation
  • Summary: The genome-wide distributions of simple tandem repeat lengths (DRLs) are surprisingly similar across mammals, suggesting they have evolved in steady state for hundreds of millions of years. Qualitative features of the DRLs are similar for different repeat motifs: the distribution rapidly decays by many orders of magnitude but transitions to a prominent tail of long repeats that truncates just below disease-associated lengths (L ~ 100 nt), despite the risk of rapidly expanding to severe neurodegenerative disorders (e.g., Huntington’s chorea). While this naïvely suggests truncation due to purifying selection, the same distribution profile is more parsimoniously explained by a selectively neutral dynamic balance between competing mutational processes: a long tail of intermediate to long repeats is a direct consequence of the non-linear increase in expansion and contraction rates referred to as ‘repeat length instability’. The neutrality of the broader genomic distribution can be explained, in part, by the large intronic and intergenic targets, which is sharply contrasts DRLs assembled from protein-coding regions that comprise the human exome. Although the exome amounts to only 1-2% of the genomic target, coding repeats are severely depleted of longer lengths primarily due to the aversion to frameshift mutations. The remaining roughly 45-50% of the human genome is comprised of various types of transposable elements (TEs), the majority of which are inactive today. Repeat lengths in this half of the genome tell a more complex story: DRLs in active TEs appear to be shaped by functional constraints, while inactive TEs, when stratified by the time of inactivation, interpolate between the active DRLs and the neutral steady-state (i.e., genome-wide DRL). Focusing on mononucleotide-A repeats, active protein-coding TEs like LINEs are depleted of long repeats, much like the exome, while active Alus harbor enriched tails associated with the A-rich terminal sequence required for retrotransposition. The phylogeny of LINE subfamilies shows a highly asymmetric progression: mass inactivation of a dominant subfamily followed by subsequent establishment of a new dominant subfamily, a repeating cycle dating back to mammalian origins. This allowed us to construct a dated series of subfamily-specific DRLs with subsequently longer time since inactivation equating to successively longer periods of neutral mutation accumulation. Assuming functional constraints on repeat lengths remained similar, this can be interpreted as a temporal trajectory of the response of the DRL to subfamily inactivation, i.e., the instantaneous relaxation of constraints. Due to the associated timescale, the non-equilibrium response of the DRL to pseudofunctionalization provides information-rich data primarily informed by repeat instability rates. We used this to infer length-dependent rates of expansion and contraction for A repeats, finding rates consistent with both 1) de novo mutation rate estimates for a subset of lengths and 2) constraints on the parameter space inferred from the shape of the whole-genome DRL under the assumption of neutral steady state evolution.

1:40PM - 7. Alex Diaz-Papkovich: Ramachandran Lab, Data Science Institute  |  Brown 

  • Title:  From Wikipedia to AI: Measuring 25 years of synthesis of human genetics research in the information ecosystem
  • Summary:  Genetics research frequently intersects with ethnicity, nationality, and race, making it uniquely vulnerable to misrepresentation. Yet, 25 years after the initial sequencing of the human genome, there is little understanding of how human genetics research lives in the public-facing information ecosystem. We analyze text and search for terminology related to genetics research from 3,050,422 historical revisions across 6,738 Wikipedia pages about ethnicity, nationality, and race spanning 25 years, as well as 56,908 user discussions from these pages. We also analyze LLM chatbot responses to queries about nationalities, and 133 pages from Grokipedia, an AI-generated encyclopedia created by Elon Musk’s company xAI. We find that in our Wikipedia pages, genetics terminology is present on 14.8% of pages (including 55.5% of our 1,000 most viewed pages), and in 67.8% of pages about nationalities. It is often presented in the form of separate “Genetics” sections, suggesting research is synthesized to present ethnicity and nationality as having a biological element. LLM responses mention genetics in 9.5% to 31.4% of responses (depending on the LLM) and cite Wikipedia in 82.5% to 100% of responses, suggesting Wikipedia has a downstream influence on chatbots. Lastly, we find Grokipedia mentions genetics more frequently than Wikipedia in the context of nationality, that it explicitly defines racial groups by their genetics and asserts genetics is causal for social outcomes such as education and crime, and that it frequently hallucinates or misrepresents scientific sources. Taken together, these results suggest human genetics research has been synthesized over time into a genetic essentialist representation of group identity, and that such a framing can be accelerated by AI technologies.


2:00PM -  Closing Remarks [Organizers]

**Due to unforeseen circumstances, the schedule may change.