Legal
Data ethics
A platform that simulates people can be built responsibly or irresponsibly, and the difference is mostly in constraints accepted before the product exists. These are ours, including where they cost us capability.
Aggregation is architectural, not cosmetic
The k≥100 floor is enforced when cohorts are constructed. Groupings below it are discarded, not retained-and-hidden. There is no configuration flag, no enterprise tier, and no support request that produces a cohort of one. The product cannot simulate a named individual because the pipeline never builds one.
This costs us something real: the most commercially attractive version of this technology is a conversational twin of a specific creator, and several competitors sell exactly that. We do not, and we think that is the right call.
Re-identification
Cohort packs contain aggregate counts and unattributed verbatim text. Verbatim quotes are the highest-risk element — a distinctive enough sentence can be searched for. We accept that risk knowingly because retrieved voice is what keeps a cohort from answering like a stereotype, and we mitigate it by never attaching a handle, a link, or a timestamp to a quote.
Attempting re-identification is grounds for immediate termination. That is the one term we will not negotiate.
The consent problem, stated honestly
The behavioural data is publicly visible. That is a legal basis; it is not the same as informed consent, and we will not pretend otherwise. Someone who posted a video in 2023 did not agree to be one of 369 accounts inside a cohort that informs an advertising decision in 2026.
What we do about it: aggregate so no individual is the subject of a result, never let account-level data reach a model or an export, honour exclusion requests by rebuilding the affected aggregate, and decline use cases that target individuals. What we cannot do is claim consent we do not have.
Uses we decline
- Profiling or targeting a specific named individual.
- Decisions with legal or similarly significant effects on people — credit, employment, housing, insurance, benefits.
- Inferring protected characteristics — health, sexuality, religion, immigration status — from behavioural signal.
- Political microtargeting, and any use designed to depress participation or manufacture the appearance of grassroots consensus.
- Passing off simulated output as statements by real, identified people.
Honest reporting
Every reported share carries an interval. Every evaluation reports the no-persona and trivial baselines alongside the result. Cohorts with thin evidence are labelled thin on the cohort itself rather than in a footnote. When a method fails to beat a demographic baseline, that is published in the same place and at the same prominence as when it succeeds.
The reason is not modesty. A simulation platform whose failures are invisible is indistinguishable from one that does not work, and the only way out of that is to make the failures visible on purpose.