Functional Genomic Representation Learning for Microbial Communities
Jan 1, 2026
·
1 min read

Proteogenic k-mer tokenization yields interpretable, parameter-free genome encodings that plug into both classical statistics and pretrained sequence models (Evo, ESM, ProtT5). Abundance-weighted community embeddings, leakage-safe features, permutation tests and honest cross-validation.

Authors
B.S.E. Candidate, Data Science
I study Data Science at the University of Michigan, Ann Arbor. I came here through the UM-SJTU dual-degree program: the first half of my undergraduate degree was Mechanical Engineering at Shanghai Jiao Tong University’s Joint Institute, the second half is Data Science at Michigan. It was less a change of heart than a change of tools.