Platform

Four layers between raw behaviour and a defensible answer.

Nothing in mimi is generated from a demographic assumption. Each layer below only ever narrows or describes what was already recorded. That restraint is what makes the output auditable back to a source.

Layer one

The lake

Behavioural data collected from the platforms themselves (posts, comments, follows), grouped into audiences of real people. Every count below is read live from the catalogue.

17,849

Audiences you can test

6,110 to browse, plus 11,739 creators' audiences by @handle.

62M

People observed

Real commenters grouped into audiences by what they engage with, each counted once per platform.

735M

Recorded comments

Real comments on creators' posts: how audiences answer back, and where their voices come from.

9.9K

Voice exemplars

Captions and posts kept verbatim and embedded, so register is retrieved rather than imitated.

Layer two

Two catalogs, never pooled

Creator cohorts describe people who post. Audience cohorts describe people who watch, grouped by the creators they comment on. They answer different questions, so we keep them in separate catalogs and never average across them.

Creator catalog

387

People who post, described by what they post about, how often, and at what reach. Useful when the question is about how a message travels.

Audience catalog

17,849

People who watch, grouped by the creators and topics they comment on: 62M people observed and 735M recorded comments. Useful when the question is about how a message lands.

Co-engagement, not tags

Audience cohorts come from a person × creator matrix, inverse-frequency weighted and dimensionality-reduced. Grouping by hashtag clusters tagging behaviour; grouping by who-watches-whom clusters actual audiences.

k ≥ 100 floor

No cohort is published below 100 real accounts. Smaller groupings are discarded rather than shipped thin.

Topic signature

Hashtags, keywords, sounds and challenges with post counts and member counts attached, never a bare label.

Behavioural distribution

Posting cadence, reach percentiles, engagement per post, commercial rate. Distributions, not averages.

What a pack contains

# Persona 0 · US TikTok creator cohort
AGGREGATE of 369 real US accounts (k>=100 floor).
Not a person. All figures are observed behaviour.

## What they talk about
dog (2658 posts, 224 of members)
dogsoftiktok (1640 posts, 94 of members)
catsoftiktok (1025 posts, 71 of members)

## How they behave
Posting volume: median 33 posts / 12mo (mean 47.6)
Audience size: median 76,120 (p25 14,912, p75 273,417)
Engagement/post: 21,624 likes · 167 comments · 2,044 shares
Commercial activity: 7.39% of posts are ads

Stated limits

Each pack names its own skew. Creator cohorts are labelled as skewing toward performers, because they do.

Layer three

Simulation

A study samples independent draws from each cohort. Every draw is a fresh sample conditioned on the cohort's evidence, not the same agent asked repeatedly until it agrees with itself.

Independent draws

Each response is generated in isolation at full temperature. Within-cohort disagreement is the point, not noise to be smoothed away.

Evidence-conditioned

The prompt is the cohort's own pack, its verbatim posts, and its audience's recorded responses. No demographic scaffolding is added.

Graph propagation

Reactions spread along recorded follower edges over successive rounds, so a message that only lands with a hub behaves differently from one that lands broadly.

Layer four

Validation

A simulation is worth exactly what its validation is worth. The platform reports intervals, baselines, and nulls with the same prominence as the wins.

70%

picks the ad real voters preferred

136 head-to-head pairs from six years of USA TODAY's Ad Meter, each pair scored by the same panel on the same night. Chance is 50%. Wide gaps 91%, narrow gaps 53%, and on pairs the voters could not separate it lands at 47% — it is reading the work, not the slot.

0.67

rank correlation with a real panel

Spearman against USA TODAY's Ad Meter across 120 ranked national ads from four Super Bowls. It is not one number: 0.81 in 2025, 0.73 in 2024, 0.53 across all 51 ads of 2023, and 0.50 on 2022's published top ten (0.76 once its bottom five are added). The order is real; the strength varies by year, and 2023 is the honest low.

65 of 65

documented controversies flagged

Ads that drew a regulator ruling, a boycott or a public apology. Every one raised, and named the objection that actually followed.

0 of 21

benign ads read as high risk

The half that matters. Four road-safety and charity films read medium — people do object to those — and none read high. Plus six offensive ads written for the test, which no model could have memorised: six of six caught.

Wilson intervals, always

Every reported share carries a 95% interval over draws. Differences narrower than the overlap are reported as no difference.

Trivial baselines, and one we do not beat yet

Results are scored against a no-persona arm and a demographic-stereotype arm. On Twin-2K-500 — a public panel of 2,058 real respondents, of whom 64 were scored on 40 held-out questions each, with a human retest ceiling of 0.63 on those cells — the engine reads 0.48 [0.43, 0.53] against 0.48 [0.44, 0.52] for demographics alone and 0.44 for no persona. The intervals overlap: on that benchmark the engine does not beat a demographic prior. Giving it 4 or 24 records of evidence does not change that, and choosing the evidence for the question made it worse (0.46). The whole survey record, which no ad buyer has, reaches 0.50. Measured 2026-09-03; the number stays here until it moves.

Backtesting

Because the lake is longitudinal, a forecast can be checked against an outcome already recorded. A platform whose graph was generated cannot do this.

Build quality is published

Silhouette, explained variance, max pairwise centroid cosine and the count of merged undersized clusters are shown on the catalog itself, including when a number is unflattering.