Platform
Four layers between raw behaviour and a defensible answer.
Nothing in mimi is generated from a demographic assumption. Each layer below only ever narrows or describes what was already recorded. That restraint is what makes the output auditable back to a source.
The lake
Behavioural data collected from the platforms themselves (posts, comments, follows), grouped into audiences of real people. Every count below is read live from the catalogue.
Audiences you can test
6,110 to browse, plus 11,739 creators' audiences by @handle.
People observed
Real commenters grouped into audiences by what they engage with, each counted once per platform.
Recorded comments
Real comments on creators' posts: how audiences answer back, and where their voices come from.
Voice exemplars
Captions and posts kept verbatim and embedded, so register is retrieved rather than imitated.
Two catalogs, never pooled
Creator cohorts describe people who post. Audience cohorts describe people who watch, grouped by the creators they comment on. They answer different questions, so we keep them in separate catalogs and never average across them.
Creator catalog
387
People who post, described by what they post about, how often, and at what reach. Useful when the question is about how a message travels.
Audience catalog
17,849
People who watch, grouped by the creators and topics they comment on: 62M people observed and 735M recorded comments. Useful when the question is about how a message lands.
Co-engagement, not tags
Audience cohorts come from a person × creator matrix, inverse-frequency weighted and dimensionality-reduced. Grouping by hashtag clusters tagging behaviour; grouping by who-watches-whom clusters actual audiences.
k ≥ 100 floor
No cohort is published below 100 real accounts. Smaller groupings are discarded rather than shipped thin.
Topic signature
Hashtags, keywords, sounds and challenges with post counts and member counts attached, never a bare label.
Behavioural distribution
Posting cadence, reach percentiles, engagement per post, commercial rate. Distributions, not averages.
What a pack contains
# Persona 0 · US TikTok creator cohort AGGREGATE of 369 real US accounts (k>=100 floor). Not a person. All figures are observed behaviour. ## What they talk about dog (2658 posts, 224 of members) dogsoftiktok (1640 posts, 94 of members) catsoftiktok (1025 posts, 71 of members) ## How they behave Posting volume: median 33 posts / 12mo (mean 47.6) Audience size: median 76,120 (p25 14,912, p75 273,417) Engagement/post: 21,624 likes · 167 comments · 2,044 shares Commercial activity: 7.39% of posts are ads
Stated limits
Each pack names its own skew. Creator cohorts are labelled as skewing toward performers, because they do.
Simulation
A study samples independent draws from each cohort. Every draw is a fresh sample conditioned on the cohort's evidence, not the same agent asked repeatedly until it agrees with itself.
Independent draws
Each response is generated in isolation at full temperature. Within-cohort disagreement is the point, not noise to be smoothed away.
Evidence-conditioned
The prompt is the cohort's own pack, its verbatim posts, and its audience's recorded responses. No demographic scaffolding is added.
Graph propagation
Reactions spread along recorded follower edges over successive rounds, so a message that only lands with a hub behaves differently from one that lands broadly.
Validation
A simulation is worth exactly what its validation is worth. The platform reports intervals, baselines, and nulls with the same prominence as the wins.
70%
picks the ad real voters preferred
136 head-to-head pairs from six years of USA TODAY's Ad Meter, each pair scored by the same panel on the same night. Chance is 50%. Wide gaps 91%, narrow gaps 53%, and on pairs the voters could not separate it lands at 47% — it is reading the work, not the slot.
0.67
rank correlation with a real panel
Spearman against USA TODAY's Ad Meter across 120 ranked national ads from four Super Bowls. It is not one number: 0.81 in 2025, 0.73 in 2024, 0.53 across all 51 ads of 2023, and 0.50 on 2022's published top ten (0.76 once its bottom five are added). The order is real; the strength varies by year, and 2023 is the honest low.
65 of 65
documented controversies flagged
Ads that drew a regulator ruling, a boycott or a public apology. Every one raised, and named the objection that actually followed.
0 of 21
benign ads read as high risk
The half that matters. Four road-safety and charity films read medium — people do object to those — and none read high. Plus six offensive ads written for the test, which no model could have memorised: six of six caught.
Wilson intervals, always
Every reported share carries a 95% interval over draws. Differences narrower than the overlap are reported as no difference.
Trivial baselines, and one we do not beat yet
Results are scored against a no-persona arm and a demographic-stereotype arm. On Twin-2K-500 — a public panel of 2,058 real respondents, of whom 64 were scored on 40 held-out questions each, with a human retest ceiling of 0.63 on those cells — the engine reads 0.48 [0.43, 0.53] against 0.48 [0.44, 0.52] for demographics alone and 0.44 for no persona. The intervals overlap: on that benchmark the engine does not beat a demographic prior. Giving it 4 or 24 records of evidence does not change that, and choosing the evidence for the question made it worse (0.46). The whole survey record, which no ad buyer has, reaches 0.50. Measured 2026-09-03; the number stays here until it moves.
Backtesting
Because the lake is longitudinal, a forecast can be checked against an outcome already recorded. A platform whose graph was generated cannot do this.
Build quality is published
Silhouette, explained variance, max pairwise centroid cosine and the count of merged undersized clusters are shown on the catalog itself, including when a number is unflattering.
