Project SIMS

Platform

Four layers between raw behaviour and a defensible answer.

Nothing in Project SIMS is generated from a demographic assumption. Each layer below only ever narrows or describes what was already recorded — which is what makes the output auditable back to a source.

Layer one

The lake

Longitudinal behavioural data collected from the platforms themselves — posts, engagement, follows, comments — resolved into cross-platform identities.

3.3M

Resolved identities

One person, several handles. TikTok, Instagram, YouTube and X linked where the evidence supports it.

197B

Follower edges

Recorded relationships, not a generated graph. This is the substrate opinion travels along.

311K

Audience reactions

Real comments on cohort members' content — how their audience actually answers back.

9.9K

Voice exemplars

Captions and posts kept verbatim and embedded, so register is retrieved rather than imitated.

Layer two

Two catalogs, never pooled

Creator cohorts describe people who post. Audience cohorts describe people who watch — grouped by which creators they actually comment on. They answer different questions, so we keep them in separate catalogs and never average across them.

Creator catalog

485

People who post, described by what they post about, how often, and at what reach. Useful when the question is about how a message travels.

Audience catalog

652

People who watch, grouped from 8.4M recorded comment relationships across 853K people and 145K creators. Useful when the question is about how a message lands.

Co-engagement, not tags

Audience cohorts come from a person × creator matrix, inverse-frequency weighted and dimensionality-reduced. Grouping by hashtag clusters tagging behaviour; grouping by who-watches-whom clusters actual audiences.

k ≥ 100 floor

No cohort is published below 100 real accounts. Smaller groupings are discarded rather than shipped thin.

Topic signature

Hashtags, keywords, sounds and challenges with post counts and member counts attached — never a bare label.

Behavioural distribution

Posting cadence, reach percentiles, engagement per post, commercial rate. Distributions, not averages.

Stated limits

Each pack names its own skew. Creator cohorts are labelled as skewing toward performers, because they do.

What a pack actually contains

# Persona 0 — US TikTok creator cohort
AGGREGATE of 369 real US accounts (k>=100 floor).
Not a person. All figures are observed behaviour.

## What they talk about
dog (2658 posts, 224 of members)
dogsoftiktok (1640 posts, 94 of members)
catsoftiktok (1025 posts, 71 of members)

## How they behave
Posting volume: median 33 posts / 12mo (mean 47.6)
Audience size: median 76,120 (p25 14,912, p75 273,417)
Engagement/post: 21,624 likes · 167 comments · 2,044 shares
Commercial activity: 7.39% of posts are ads

Layer three

Simulation

A study samples independent draws from each cohort. Every draw is a fresh sample conditioned on the cohort's evidence — not the same agent asked repeatedly until it agrees with itself.

Independent draws

Each response is generated in isolation at full temperature. Within-cohort disagreement is the point, not noise to be smoothed away.

Evidence-conditioned

The prompt is the cohort's own pack, its verbatim posts, and how its audience actually responds. No demographic scaffolding is added.

Graph propagation

Reactions spread along recorded follower edges over successive rounds, so a message that only lands with a hub behaves differently from one that lands broadly.

Layer four

Validation

A simulation is worth exactly what its validation is worth — so the platform reports intervals, baselines, and nulls with the same prominence as the wins.

Wilson intervals, always

Every reported share carries a 95% interval over draws. Differences narrower than the overlap are reported as no difference.

Trivial baselines

Results are scored against a no-persona arm and a majority-class arm. A model that cannot beat those has not earned the word 'prediction'.

Backtesting

Because the lake is longitudinal, a forecast can be checked against an outcome already recorded — structurally unavailable to a platform whose graph was generated.

Build quality is published

Silhouette, explained variance, max pairwise centroid cosine and the count of merged undersized clusters are shown on the catalog itself — including when a number is unflattering.