Project SIMS

About

We built the data first and the simulation second — which is the opposite of how this usually goes.

Project SIMS started as an infrastructure problem, not a product idea. Years of longitudinal behavioural data existed across TikTok, Instagram, YouTube and X before anyone asked what could be simulated with it. That ordering is the whole difference.

Position

The claim we are actually making

A synthetic audience is easy to build and almost impossible to check. Generate a plausible network, populate it with personas inferred from a demographic profile, and you have something that produces confident answers immediately. Whether those answers correspond to anything is a separate question that the architecture makes very hard to ask.

Our claim is narrower and, we think, more useful: the cohorts in this platform are aggregates of real accounts, the graph between them was recorded rather than generated, and because both are longitudinal, a prediction can be scored against an outcome that already happened. That last property is the one that cannot be retrofitted.

What we are not claiming

We are not claiming to model the general population. The lake is creator-adjacent by construction, and cohorts inherit that skew. We are not claiming simulation replaces research — it narrows what you take into research. And we are not claiming a headline accuracy figure, because a single number across all question types would be meaningless even if it were true.

Principles

What the product is built against

Aggregates, never individuals

A k≥100 floor is enforced at cohort construction, not filtered at display time. There is no configuration that produces a cohort of one, and no product surface that simulates a named person.

The interval is part of the answer

A share without a band is not a result. Where the evidence is thin, the platform says so on the cohort rather than quietly returning a confident number.

Read-only against the source

The application holds read credentials against every upstream system. No code path in the platform writes to, deletes from, or reindexes the underlying lake.

Publish the nulls

Evaluations that fail to beat a demographic baseline are reported in the same place, at the same prominence, as the ones that succeed.

Provenance

Where the data comes from

The lake is assembled from publicly visible activity on TikTok, Instagram, YouTube and X: posts and their engagement counts, follower relationships, public comments, and video transcripts. Accounts are resolved across platforms only where the evidence supports the link.

Nothing in the platform ingests private messages, purchased contact lists, or data behind a login the account holder did not make public. Individual account records are retained for cohort auditing and are visible only to operators — they are excluded from every cohort brief, every simulation, and every export.