Mingling Ge, MS
Research Coordinator

Research Coordinator

Mingling Ge, MS, is a Research Coordinator in the Wagenaar Lab at the Center for Neuroengineering & Therapeutics, where he builds the data infrastructure behind Epilepsy.Science – an NIH-funded initiative to make multimodal human epilepsy data discoverable and reusable across institutions. His work bridges biostatistics and biomedical informatics: he designs metadata models and schemas for intracranial and scalp EEG datasets on the Pennsieve platform, migrates NINDS Common Data Elements into queryable clinical registries, and develops the harmonization standards that let separately-collected epilepsy cohorts be compared and combined.
He holds an MS in Quantitative Biomedical Sciences (Health Data Science) from Dartmouth, where his research spanned latent-variable and mixed-effects modeling of longitudinal pediatric data and multimodal physiological signal processing. He is first author of a 2025 study on glycemic variability in adolescents in Current Developments in Nutrition.
His broader interest is in the measurement and data-quality problems that emerge when biomedical data are pooled across sites and instruments – how to make heterogeneous clinical and neurophysiological data trustworthy enough for rigorous downstream analysis. That interest now extends to machine-generated evidence: he is developing methods to verify the trustworthiness of AI-produced scientific claims against their cited sources.
Email: allen.ge@pennmedicine.upenn.edu
🎓 Google Scholar
🔗 GitHub
Overview: As part of this NIH-funded initiative, he builds the metadata infrastructure that lets intracranial-EEG (iEEG) datasets collected at different institutions be discovered, compared, and combined. He is the primary drafter of the Epilepsy.Science Data Structure (ESDS v0.6), a submission specification for harmonized BIDS-iEEG data sharing, and he standardized the CNT epilepsy metadata into five Pennsieve data models. He curated and verified 200+ BIDS-aligned iEEG sidecars (1,042 electrodes) supporting the first public release of 36 CNT epilepsy datasets.
Overview: He migrated the full NINDS Epilepsy Common Data Element catalog (875 elements across 37 CRF templates) into a queryable Pennsieve registry with a 30-property metadata model and source-to-record verification, and developed REDCap evaluation instruments from 146 CDEs for NIH expert review. To validate provenance, he reverse-engineered the NLM CDE Repository API to assemble a 2,255-record reference corpus and audit 330 live clinical CDEs, resolving cross-steward identifier collisions and retired-record masking.
Overview: He is developing a trust-verification layer for an AI-driven literature-enrichment pipeline (retrieval-augmented generation) that surfaces source-cited facts onto dataset pages. The framework combines a deterministic quote-verification step, model-based entailment and evidence-strength assessment, and a metadata-consistency guard that cross-checks generated claims against structured record tables — a data-quality approach for evidence produced by language models.
Overview: He designed a dual-key (person × session) longitudinal metadata architecture for the NIH PREVeNT neonatal-encephalopathy trial and deployed it across 80 Pennsieve datasets.