Simple and Effective Embedding Model for Single-cell Biology Built from ChatGPT
Overview
Authors
Affiliations
Large-scale gene-expression data are being leveraged to pretrain models that implicitly learn gene and cellular functions. However, such models require extensive data curation and training. Here we explore a much simpler alternative: leveraging ChatGPT embeddings of genes based on the literature. We used GPT-3.5 to generate gene embeddings from text descriptions of individual genes and to then generate single-cell embeddings by averaging the gene embeddings weighted by each gene's expression level. We also created a sentence embedding for each cell by using only the gene names ordered by their expression level. On many downstream tasks used to evaluate pretrained single-cell embedding models-particularly, tasks of gene-property and cell-type classifications-our model, which we named GenePT, achieved comparable or better performance than models pretrained from gene-expression profiles of millions of cells. GenePT shows that large-language-model embeddings of the literature provide a simple and effective path to encoding single-cell biological knowledge.
Small, Open-Source Text-Embedding Models as Substitutes to OpenAI Models for Gene Analysis.
Gan D, Li J bioRxiv. 2025; .
PMID: 40027770 PMC: 11870524. DOI: 10.1101/2025.02.15.638462.
EpiFoundation: A Foundation Model for Single-Cell ATAC-seq via Peak-to-Gene Alignment.
Wu J, Wan C, Ji Z, Zhou Y, Hou W bioRxiv. 2025; .
PMID: 39975086 PMC: 11839112. DOI: 10.1101/2025.02.05.636688.
Benchmarking large language models for genomic knowledge with GeneTuring.
Hou W, Shang X, Ji Z bioRxiv. 2023; .
PMID: 36993670 PMC: 10054955. DOI: 10.1101/2023.03.11.532238.