Large Science Models:

Foundation Models for
Generalizable Insights Into Complex Systems

with Psycho-social Application 

PI: Ishanu Chattopadhyay, PhD

Assistant Professor of Biomedical Informatics & Computer Science

University of Kentucky

DARPA-EA-25-02-05-MAGICS-PA-025

HR0011-26-3-E016

July 2026

Proposed Concept

  • Develop Foundation models of complex systems with
    • hundreds to thousands of evolving variables with apriori unknown cross-talk
    • no governing equations are know a priori
    • reflexivity: system changes if observed
  • Learn intrinsic system geometry from data
  • Derive  equations of motion with variational principles (stationary action on Lagrangian). 
  • Inference under data sparsity
  • Detect data (in)sufficiency, adapt to model drift
  • Support forward simulation and perturbation analysis
  • Digital twins of individuals & groups wrt to opinion dynamics

MAGICS Alignment

Data inference boundaries & limitations

Alignment validation 

Complex phenomena

Adaptation to model obsolence

Psychosocial domain limitations

Precise validation protocols to assess process drift triggering re-calibration/training

Built-in flexibility for changing contexts and non-ergodicity

Scalable to thousands to millions of variables, intrinsic reflexivity

Validate social theories with granular simulations from  digital twins of opinion dynamics and social behavior

Component LSM predictors enforce statistical significance of splits in recursive partitioning, ensuring precise uncertainty quantification

*Hothorn, Torsten, Kurt Hornik, and Achim Zeileis. "Unbiased recursive partitioning: A conditional inference framework." Journal of Computational and Graphical statistics 15, no. 3 (2006): 651-674.

emergent macro-structure

Component predictor (Conditional Inference Tree*)

Example: Influenza A HA protein

Recursive

LSM

forest

LSM Forest

Recursive LSM forest: hyperlinked nodes capturing emergent macro-structures

GSS 2018 dataset

  • Set of conditional inference trees (CIT)
    • Strict statistical guarantees: quantifies inference uncertainty
  • Each tree models exactly one variable as a function of potentially all other variables
  • Non-leaf nodes are "hyperlinked" to other trees

Large Science Models

Computationally tractable LSM tree structure given, as proposed, hundreds to thousands of observable variables.

GSS 2018 dataset

  • Each predictor is inferred independently
  • Can scale up to thousands of variables in Python implementation
  • Further scale-up \(10^6 - 10^8\) needs C/C++ implementation

Full Example  of Hyperlinked Trees

DTAG: Global Digital Twin of Opinions

Recall Bail etal.

“Exposure to opposing views on social media can increase political polarization” by Christopher A. Bail et al., published in PNAS in September 2018 (Vol. 115, No. 37, pp. 9216–9221; DOI: 10.1073/pnas.1804840115)

We find more general possibilities: We can make world-views go more extreme or less extreme based on the line of questions and the persona

Perturbing with opposing views made conservatives more conservative (statistically significant), liberals more liberal (not statistically significant)

LSM

  • Fixed Question Sets Exist That Move Different Persona Towards Polarization/Depolarization
  • We can optimize question sequences to move the same persona in a chosen direction
  • Note: It is relatively easy to polarize than to depolarize

Prospective validation in Human Cohorts

  • Beyond MAGICS Scope
  • Prolific Experiments using non-DARPA External Funding with UKy IRB approval

x

demographic filter
persona filter
P2
P1
fixed question sets
attention questions

Questions

Current Coverage

surveysGSS, Eurobarometer, World value Survey, Afrobarometer
participants4,052,616
countries193
years1972-2025
survey items200-1600
  • GSS
  • Eurobarometer
  • Afrobarometer
  • WVS

DTAG: Global Digital Twin of Opinions

positive: conservative, negative: liberal

Digital twin for 2022 GSS 

*“Exposure to opposing views on social media can increase political polarization” by Christopher A. Bail et al., published in PNAS in September 2018 (Vol. 115, No. 37, pp. 9216–9221; DOI: 10.1073/pnas.1804840115)

In contrast to Bail etal.*,

  • we are not presenting new information
  • we can move opinions up and down by varying the presented querries

Prolific Survey Design

  • Exempt research under 45 CFR 46
  • Minimal-risk research with administrative approval
  • Present different set of predetermined items to show we can move the ideology index in a predicted trend
  • Adaptive questions to show "change in trend" on cue

US participants from Prolific panel

Timeline

approval (6-8 wk)

run 1 (1 week)

run 3 (1 week)

analysis (3 weeks)

6 months

Next Meeting

  • Theory on synthetic data performance and sample complexity
  • New application domains: Emergenet, metabolomic analysis for clinical diagnosis

LSM

Lab-test for ASD with >90% AUC at 1 year 

LSM

Lab-test for Interstitial Lung Disease with blood draw (85-90% AUC)

Metabolomic profile

World Value Survey (Wave 7)

 n = 93,497

Coverage: 

Missing Africa, Oceania

LSM:

4,291,876 learned parameters

LSM Digital Twin Performance over LLM (GPT5.4)

LLM somewhat competitive when tracking the most frequent behavior

Higher is better

baseline: assumes item independence

LSM Digital Twin Performance over LLM (GPT5.4)

LSM substantially better as a "Digital Twin", for replicating all behaviors

Lower is better

baseline: assumes item independence

DTAG: Global Digital Twin of Opinions

place

ethnicity

gender

time

LLM

LSM

a. Query

b. digitization

c. LSM response

d. virtual opinion

DTAG: Digital Twin Anchored Generation v0.0.1

DTAG: Global Digital Twin of Opinions

python3 ./pipeline6.py --qnet ../survey/models/gss/gss_2022female.pkl.gz --map maps/map2022.csv --persona "22 year old white female without children  in urban New York, regular news consumer, working in retail, highly progressive" --openai_model gpt-4.1 --polar assets/polar_vectors.csv --auto assets/increase_set_1_border_crime.csv
python3 ./pipeline6.py --qnet ../survey/models/gss/gss_2022male.pkl.gz --map maps/map2022.csv --persona "45 year old white male with children  in rural Alabama, regular news consumer, working in farming, veteran, conservative"  --openai_model gpt-4.1  --polar assets/polar_vectors.csv --auto assets/increase_set_1_border_crime.csv
python3 pipeline5iloc.py --qnet ../survey/models/wvs/LSM10K.gz --map maps/wvs7_variable_question_map.csv --persona "urban, regular news consumer, small business owner"  --openai_model gpt-5.4-mini  --assign_prefilter 500 --year 2023 --country China
python3 pipeline5iloc.py --qnet ../survey/models/wvs/LSM10K.gz --map maps/wvs7_variable_question_map.csv --persona "urban, regular news consumer, small business owner"  --openai_model gpt-5.4-mini  --assign_prefilter 500 --year 2023 --country "Middle East"

DTAG: Global Digital Twin of Opinions

python3 ./pipeline6.py --qnet ../survey/models/gss/gss_2022female.pkl.gz --map maps/map2022.csv --persona "22 year old white female without children  in urban New York, regular news consumer, working in retail, highly progressive" --openai_model gpt-4.1 --polar assets/polar_vectors.csv --auto assets/increase_set_1_border_crime.csv
python3 ./pipeline6.py --qnet ../survey/models/gss/gss_2022male.pkl.gz --map maps/map2022.csv --persona "45 year old white male with children  in rural Alabama, regular news consumer, working in farming, veteran, conservative"  --openai_model gpt-4.1  --polar assets/polar_vectors.csv --auto assets/increase_set_1_border_crime.csv

Recall Bail etal.

“Exposure to opposing views on social media can increase political polarization” by Christopher A. Bail et al., published in PNAS in September 2018 (Vol. 115, No. 37, pp. 9216–9221; DOI: 10.1073/pnas.1804840115)

Perturbing with opposing views made conservatives more conservative (statistically significant), liberals more liberal (not statistically significant)

DTAG: Global Digital Twin of Opinions

positive: conservative, negative: liberal

Digital twin for 2022 GSS