SELIONE
Research

The protocol

Last updated October 10, 2026.

Open document · revised as the work advances

This page is the technical companion to the method. It documents what Selione does to a life, in order; what each number a traveler sees measures; how we test whether an Echo resembles its person rather than asserting it; what the published research nearby says; where our reasoning could be wrong; and the studies we would like to run with independent researchers. It is written to be argued with. Every curve below is a public statistic, computed and not drawn.

  1. The object of study
  2. How each layer is gathered
  3. The instruments in detail
  4. The quantities, and how they behave
  5. Measuring resemblance
  6. The Envelope: prospective accuracy
  7. Where the published work sits
  8. What is published, and what is not
  9. Threats to validity
  10. Studies we want to run
  11. What would falsify our claims
  12. What we do not claim
  13. Open questions
  14. Data, consent and partnerships
  15. References

1. The object of study

Selione builds, for one person, a model made of five layers. They are gathered differently and are never confused with each other, because each one fails in a different way.

  • The archive — dated, sourced memories, each marked said or inferred, each attached to one of eight domains of life (childhood, family, love, friendship, work, beliefs, fears, dreams) and to one of seven strata of the self (language, emotional memory, personality, choices, beliefs, life story, reflexivity). This is testimony. Its failure mode is omission.
  • The narrative — the Mirror and the Constitution: how the person organises their own life into chapters, values and turning points. This is interpretation, proposed by the system and rewritten by the person. Its failure mode is flattery.
  • The kernel — what happens in them when something they have never seen arrives. Gathered by provocation, never by introspection. Its failure mode is over-generalisation.
  • The present — an inner state recomputed nightly: preoccupation, mood, recurring people, open threads. Its failure mode is recency bias.
  • The register — sealed, dated predictions about the next month, opened and scored a month later. It is the only layer that cannot be produced after the fact, and the one the others are tested against. Its failure mode is timidity: predicting only what is certain.

2. How each layer is gathered

LayerInstrumentAnti-contamination rule
ArchiveGuided interview (404 quests, 2,290 questions, unlocked over seven stages, at most 30 quest answers a day), free conversation, voice, photographs, volunteered archivesEverything not stated outright is marked inferred, with a confidence below that of anything said; inferred memories never reach the printed Book
NarrativeSynthesis proposed by the system, rewritten by the traveler, versioned, on requestEvery version is kept; the traveler's wording always wins over ours
KernelSparks: catalogue stimuli — written scenes, photographs shown without a caption, short synthesised sounds — plus stimuli generated for that person; first unedited reaction, latency recordedNever a question about their past; their life is the aim, never the content of a stimulus; latency starts when a sound ends
PresentThe Night: one pass per traveler while they are away, over what was given recently, only if material existsConsolidated facts are stored as inferred, with a lower confidence than anything said, and are deletable
RegisterThe Envelope: a handful of refutable claims about the coming month, written during a Night and sealed, once the archive holds enough to say somethingUnreadable before opening (enforced by the database); SHA-256 anchored in Bitcoin via OpenTimestamps; the archive scores it without seeing the traveler's marks

3. The instruments in detail

3.1 Quests

A quest is a short set of questions on one theme, most of them open. Closed formats exist and are deliberately few: they are cheap for the person and give the model anchors, but they are not where a person lives. After an open answer the system may ask one follow-up that picks up a specific detail just mentioned, because autobiographical memory is reconstructed from cues and a specific cue retrieves more than a general one (Conway & Pleydell-Pearce, 2000).

Quests unlock by stage, and the order follows the life review literature (Butler, 1963; Westerhof & Bohlmeijer, 2014): concrete recall before evaluation, evaluation before integration. The daily cap exists for two reasons: fatigue lowers the specificity of recall, and a life told in one weekend is a summary, not an archive.

3.2 Sparks

People are poor reporters of why they react as they do (Nisbett & Wilson, 1977), but good sources of the reaction itself when it is captured before it is edited (Ericsson & Simon, 1980). A spark therefore never asks "how do you usually feel about…". It shows a scene, a photograph, a sentence said to you, a choice, a sound, and keeps the first thing typed and how long it took. The signature derived from sparks is rewritten as reactions accumulate. It holds tendencies, conditional rules ("when X, they Y") and the places where the person contradicts themselves, which are kept rather than smoothed.

3.3 The Night

While the traveler is away, the system reads what was given recently against what was already known. It links new material to old, turns repeated episodes into general statements (only when several separate memories support them), recomputes the present, lists what is still unknown, and writes a few unprompted bets about what the person would say on topics likely to come up. This is modelled on the function attributed to sleep in systems consolidation (Diekelmann & Born, 2010) and on accounts of cognition as prediction under correction (Clark, 2013). We borrow the shape of the loop because it makes the model testable; we make no claim about mechanism.

3.4 The Craft

The interviewer improves across everyone, but only from numbers. Each exchange is reduced to a coarse record of what the interviewer did, the situation, and what followed (reply length, richness). From those counts the interviewer's rules are revised, and a revision under which richness drops does not stay. The objective is richness and return, never time spent: rules that resemble retention techniques are rejected.

3.5 Shadow trials and retests

A solicited trial is announced, and announced measurement changes behaviour. Since September, some open quest questions are also shadow trials: the Echo's guess is written before the answer is stored, and a judge compares the two afterwards. The traveler is never told which questions were shadowed and is never asked to mark them. Shadows are bounded per day and wait for a minimum archive. Separately, some solicited trials are retests of a question the traveler answered months earlier; the agreement between the two answers, scored by the same judge, is the traveler's own consistency, and it is the ceiling against which fidelity is read (Park et al., 2024, use the same yardstick). Misses are typed — fact, value, voice, person — and feed question selection.

3.6 The verbatim register

Alongside the archive of filed memories, the system keeps the traveler's own lines as said, in the language they were said in, and draws on them when the Echo speaks. The archive carries what happened; the register carries how the person says things.

3.7 The rooms as instruments

Make me sleep is not open, so nothing below is being collected today; it is written here because a research page that drops what it closed is a research page one cannot check. Make me sleep records the answers to its safety and arousal questions, the steps taken, whether the ten minutes started and ended, and one morning rating out of four. The pathway adapts to the person: after enough rated mornings, a step whose nights went measurably better for them is kept. The morning rating is a subjective proxy and is reported as one. The Return enforces its anti-suggestion rule on words: the vocabulary of every question is checked against what the traveler has named, and the share of fragments marked "I might be filling this in" is a number that is supposed to look bad; a share near zero is treated as a defect. A second walk into the same moment, recognised by similarity, runs blind and ends with a judged consistency score. Dreams are anchored to memories by similarity; the anchoring window also covers the day before and days five to seven before the telling, following the day-residue and dream-lag effects (Nielsen et al., 2004), and the traveler judges each anchor.

3.8 The loop

Every measure — a trial, a bet of the Night, a night's note, a witness test, an Envelope — is stamped with the build and the models in force when it was made, and with the arm of every running experiment. An experiment has a key, two or more arms and a share; a traveler is assigned to an arm without anyone choosing, the assignment is fixed per person, and it is independent across experiments. Measures are read per arm and per version; a change that does not move the measures is not kept. Promotion and stopping are decisions made by a person reading the numbers, with thresholds written down in advance; no promotion is automatic.

3.9 Portraits against a public yardstick

Every measure above compares the Echo with the traveler's own words, judged by a model. Once a year the traveler also answers two short public instruments — a twenty-item Big Five inventory from the public-domain pool (Donnellan et al., 2006) and a ten-item values scale in the short single-item form of Schwartz's theory (Lindeman & Verkasalo, 2005) — and the Echo answers the same items as the traveler, from the archive, before the traveler's answers are read. Both are scored by the instruments' own rules. The report is the mean absolute distance across dimensions as a share of the scale, and the correlation of the two profiles. It is the one number about an Echo that no model in the house produces, and it is bounded by the instruments' own reliabilities: with scale alphas around .65 to .75, an agreement near .8 is close to the ceiling, not a claim of more.

3.10 Revisits, confidence, and the sleep trial

Three more instruments run without a page of their own. A memory told in enough words is asked about again at growing intervals over the following months, in a question that names nothing the traveler did not; the new telling is judged against the first for agreement, and a memory that holds twice is marked as bedrock in the Echo's digest (Cepeda et al., 2006; Roediger & Karpicke, 2006). Every shadow guess now carries the Echo's own stated confidence; below a threshold it abstains, and calibration is read per band as it is for the Envelopes. And a traveler who asks for it enters an n-of-1 sleep trial: on nights with a place of their own, the night is drawn between that place and a neutral scene, and the two arms are compared on that person's own morning ratings (Lillie et al., 2011).

4. The quantities, and how they behave

4.1 Light

Light is earned by telling. It is a bounded score per answer that grows with how much the answer actually says and with how deep the quest is; length alone earns nothing, and a rich answer to a deep quest is worth several times a one-word answer to an early one.

4.2 Stages

A stage opens when one condition holds: enough light since the journey began. Each stage's threshold is further from the last than the last was from the one before. There is no time floor — until 6 October 2026 there was one, and it made eighteen months the fastest possible path whatever the effort; it was removed as artificial. What bounds the fastest path now is the daily ceiling of 30 answers, which puts Alcyone about two months away for someone who gives it every day, and about eighteen months away at a few hours a week.

4.3 Progress through the Epic

Until September 2026 progress was computed as an equal share per stage, which made the first stage — a small share of the journey's light — look like a seventh of it. A single quest displayed 2%. Progress is now the share of the total light, and it cannot pass a threshold before that stage has actually opened.

4.4 Coverage

Coverage says how much of one domain of a life has been told. It combines two things: what was said, on a curve that starts slowly and saturates without ever reaching its ceiling, and how far the person walked through that domain's quests. The slow start is deliberate: the first dozen memories about a childhood teach almost nothing, and it is the hundreds in the middle that turn fragments into a life, so a domain is only well covered after several hundred memories. A plain exponential rises fastest at zero, which gave away several percent of a whole domain for seven questions — it was corrected on 21 September 2026. The quest term rewards breadth: the same number of memories covers more when it comes from many different questions. Life coverage is the mean over the eight domains.

4.5 Kernel depth

Kernel depth rises with the number of reactions recorded: quickly at first, then more slowly, and it never reaches one. It keeps refining and never finishes.

4.6 Consistency, and relative fidelity

The retest ceiling is the mean agreement between a traveler's two answers to the same question, months apart. Fidelity is reported beside it, and relative fidelity is their ratio, capped at one: an Echo cannot be asked to agree with a person more than the person agrees with themselves.

ceiling = mean over retests of s(a₁, a₂) s ∈ [0, 1], the judge's score relative fidelity = min(1, fidelity / ceiling)

4.7 Stability of a memory

A memory revisited at growing intervals — the spacing that consolidates recall (Cepeda et al., 2006; Roediger & Karpicke, 2006) — earns a stability: the share of revisits on which the detail asked about came back the same. Bedrock is a memory that held every time it was revisited; drift is one that changed. The Echo is meant to lean on bedrock. This measure is being introduced and is reported as such.

stability(m) = (1/k) · Σ agree_i k revisits, agree_i ∈ {0, ½, 1}

5. Measuring resemblance

We never show a percentage of a person. We show measurements of resemblance on specific tasks, each designed so that the Echo cannot have seen the answer.

5.1 Solicited fidelity

The Echo answers a question the traveler has never answered. Its answer is stored before the traveler is shown the question. The traveler then answers, and only afterwards sees the Echo's attempt. A third pass scores the pair from 0 to 1 on substance and on manner, and the traveler can overrule it.

5.2 Passive fidelity

Each night the Echo writes bets about topics likely to come up. When the traveler later talks about one of them unprompted, the bet is retrieved by similarity and a judge decides whether they are about the same thing before scoring. The similarity floor is deliberately low because the judge can refuse; scoring an unrelated pair would poison the curve. Nobody is asked anything, the bets judged in a day are bounded, and the curve falls as easily as it rises.

5.3 The witness test

Someone who knows the traveler well sees pairs of answers, one from the person and one from the Echo, unlabelled, and picks the real one. Fifty percent is the meaningful target: it means they could not tell. Anything well above means the Echo is recognisably not the person. With only five rounds, chance alone produces wildly different scores, which is why single tests are never reported as evidence.

Scores in a five-round witness testProbability of each score when the witness is guessing, compared with a witness who recognises the person four times out of five. A 5/5 happens 3% of the time by pure chance.
Cannot tell them apart (p = 0.5)Recognises the person 80% of the time
0%10%20%30%40%012345Correct picks out of 5Probability

Pooled across tests, we report the proportion with a Wilson interval, which behaves well near 0.5 and at small n (Wilson, 1927):

p̂ ± half-width, half-width = z / (1 + z²/n) · √( p̂(1−p̂)/n + z²/(4n²) ), z = 1.96
How precise a pooled witness rate isMargin around 50% at 95% confidence. Below about 100 rounds, a result of 50% cannot be told apart from 40% or 60%.
0%10%20%30%100200300Pooled rounds± margin

The new-spark variant. Some rounds of a test use a stimulus the traveler reacted to after the kernel was last rewritten. The Echo answers from the signature alone and has never seen that reaction. These rounds are scored separately: they are the direct test of whether the kernel generalises to material the archive never described.

6. The Envelope: prospective accuracy

Everything above is still retrospective in a sense: the Echo answers about a life that has already happened. A language model is very good at sounding right about the past. The Envelope removes that advantage by asking the Echo to be right about the future, in writing, before the fact — a single-person version of preregistration (Nosek et al., 2018).

  • A handful of claims, spread across several kinds of thing — what will occupy them, what they will postpone, a choice they will make — never one kind only.
  • Each claim carries a stated probability, and the register is built so that it cannot consist only of safe bets.
  • Excluded by instruction and by a filter afterwards: health, death, grief, money trouble, separation, pregnancy, and anyone else's behaviour.
  • The traveler cannot read an envelope before it opens — enforced by a database policy, not by the interface — to prevent self-fulfilling predictions (Merton, 1948).
  • The exact text is hashed with SHA-256 and the hash is anchored through OpenTimestamps, so the date of writing can be verified without trusting us (Haber & Stornetta, 1991).

6.1 Calibration

Accuracy alone rewards timidity. We therefore score the stated probabilities with the Brier score, a strictly proper scoring rule: a forecaster minimises it only by stating their true belief (Brier, 1950; Gneiting & Raftery, 2007).

BS = (1/N) · Σ (f_i − o_i)² f_i stated probability, o_i = 1 if the claim held, else 0 Murphy (1973): BS = reliability − resolution + uncertainty
Expected Brier score by stated confidenceFor claims that come true at a given rate, the score is lowest exactly when the stated confidence equals that rate. Always saying 50% scores 0.25.
Claims that come true 60% of the timeClaims that come true 85% of the time
0.000.200.400.600.800.000.200.400.600.801.00Stated confidenceBrier score

6.2 Two independent verdicts

The traveler marks each line true, false or not yet (true and false are final). Separately, a pass over the month's record — memories, quest answers, messages — gives its own verdict, including "no evidence". The traveler never sees the archive's verdict before marking. On lines where the archive had evidence, we report agreement, and in research aggregates Cohen's kappa, which corrects for agreement by chance (Cohen, 1960):

κ = (p_o − p_e) / (1 − p_e) p_o observed agreement, p_e agreement expected by chance

6.3 Is it better than a base rate?

A claim like "you will keep postponing X" may be true for most people. The honest comparison is not with 50% but with the base rate of that kind of claim, estimated across all envelopes. The number of judged lines needed to show a real difference is not small: individual registers become informative only after a long run of envelopes; cohort analyses much sooner.

Judged lines needed to beat a base rateOne-sided test at α = 0.05 with 80% power. Showing 70% against a 60% base rate takes on the order of 140 lines — years of envelopes for one person, or a few weeks across a cohort.
Base rate 50%Base rate 60%
020040060065%70%75%80%85%90%True accuracyLines needed

7. Where the published work sits

The method page lists the works that shaped the questions. Here are those that bear on the machinery and on the claims.

  • Simulating individuals from interviews. Park et al. (2024) built agents for 1,052 people from two-hour interviews; the agents reproduced participants' answers to the General Social Survey 85% as accurately as the participants reproduced their own answers two weeks later. This is the closest published result to what an Echo attempts, with far less material than an Epic provides. It also shows the right yardstick: a person's own consistency, not perfection.
  • Language models as samples of people. Argyle et al. (2023) showed that conditioning a model on demographic backstories reproduces response distributions of human subgroups. It is the population-level cousin of our problem, and a warning: a model can match a group while knowing nothing of the individual.
  • Computers judging personality. Youyou, Kosinski & Stillwell (2015) found that judgments from digital footprints can exceed those of friends; Park et al. (2015) predicted personality from language. Both support the idea that text carries a person; neither shows that a generated text is the person.
  • Self versus others. Vazire & Mehl (2008) found that other people predict someone's daily behaviour about as well as that person does, and in part through different information. This is why the Envelope is judged twice, and why witnesses matter.
  • Limits of introspection. Nisbett & Wilson (1977) and Ericsson & Simon (1980) ground the kernel's refusal to ask people to explain their reactions.
  • Memory is reconstructive. Schacter (1999) and Loftus (2005) document how suggestion and repetition create false detail. It is why inferred memories are labelled, weighted less, correctable, and never printed.
  • Forecasting. Mellers et al. (2014) showed that training, teaming and tracking improve probabilistic forecasts, and that calibration is learnable. The Envelope feeds each Echo its own past misses and the anonymous calibration of all envelopes.
  • Tacit knowledge. Polanyi (1966) underlies the kernel's premise that people know more than they can say; Kahneman & Klein (2009) set out when expert intuition deserves trust.
  • Reminiscence and wellbeing. Reviews of life review and reminiscence interventions (Westerhof & Bohlmeijer, 2014) report benefits in structured settings. We make no therapeutic claim; this literature informs pacing and tone.
  • Bereavement. Continuing bonds (Klass, Silverman & Nickman, 1996) and the dual process model (Stroebe & Schut, 1999) shape how recipients meet an Echo; prolonged grief criteria (Prigerson et al., 2009) define what we must not make worse.

8. What is published, and what is not

Section 3 calls quests, sparks, the Night and the Craft instruments on purpose. Every instrument here is grounded in cited, publicly available research, and we want it checked line by line — that is the entire point of a page like this one.

Published, in full, on this page and the method page:

  • The five layers, what each is for, and how each fails.
  • What each instrument is, what it promises, and what it is forbidden to do.
  • The formulas of every public statistic — the Brier score, the Wilson interval, Cohen's kappa, relative fidelity, stability — and the floors under which nothing is published.
  • The research each design choice is grounded in, cited throughout.
  • What we test, how, and what would prove us wrong, below.
  • Every safeguard on the method page, and what triggers it.

Not published: the internal thresholds, prompts and heuristics the product runs on. They change as the work advances, and none of the promises above depends on them; the measurements are what hold us to the promises.

9. Threats to validity

ThreatHow it could fool usWhat we do about it
Performance for the machinePeople may present a curated self, so the Echo resembles the performanceSparks capture unedited first reactions; imports bring text written for other people; witnesses judge against the person they know
LeakageThe Echo sees the answer it is being tested onFidelity guesses are stored before the question is shown; new-spark rounds use reactions newer than the kernel; envelopes are written before the month
Lenient self-markingTravelers mark envelope lines true to be kind to their EchoAn independent archive verdict; true and false are final; agreement is reported, not hidden
Plausibility bias in judgesAn AI judge rewards fluent, generic answersJudges must first decide whether a pair is about the same thing; travelers can overrule; witness tests use humans
SurvivorshipOnly engaged people reach Alcyone, so late-stage results flatter the methodResults are reported by stage and by cohort, including people who stopped
Model driftChanging the underlying language model changes the Echo overnightThe archive, kernel and register are model-independent; fidelity is re-measured after any provider change
Base-rate claimsEnvelopes full of near-universal predictions look accurateKind-level base rates across all envelopes; Brier score; a register that cannot consist only of safe bets
ReactivityBeing interviewed daily changes the person being modelledMeasured, not assumed away: see study 6 below
A judge of the same houseAn AI judge rates answers from its own model family higherThe judge runs at a different provider than the Echo it scores (since 26 September 2026); every measure carries the judge's model, so a change re-measures fidelity rather than blending it
Contaminated experimentsA traveler who moves between arms blurs bothAssignment is made without anyone choosing and is fixed for life; every measure carries its arm and its version

10. Studies we want to run

None of these has been run. Each would be preregistered with its analysis plan before any data is examined, run only on travelers who opted in to research, and reported whatever the result.

Study 1 — Does time build resemblance?

Question. Does an Echo built over months resemble its person more than one built from the same number of answers in a week? Design. Compare witness rates and solicited fidelity at matched answer counts between travelers who reached them slowly and a short-format comparison group. Prediction. Slower accumulation yields higher fidelity at equal volume, because recall is cued by the days in between.

Study 2 — Does the kernel generalise?

Question. On stimuli the Echo has never seen, do witnesses confuse its reactions with the person's?Design. Within-subject comparison of new-spark rounds with archive rounds; an ablation where the Echo answers without its kernel. Prediction. Witness accuracy on new sparks moves towards 50% as kernel depth rises, and the ablation moves it away.

Study 3 — Can an Echo forecast its person?

Question. Is envelope accuracy above kind-level base rates, and does calibration improve over the year? Design. Mixed-effects logistic model of line outcomes with stated confidence, kind and month as predictors; Brier decomposition by quarter. Prediction. Resolution rises and reliability falls over time.

Study 4 — How often does the system invent people?

Question. What share of inferred memories do travelers correct or delete, and does it fall as coverage rises? Measure. Correction and deletion rates for inferred versus said memories. This is our most direct measure of confabulation.

Study 5 — Self, others, and the archive

Question. Following Vazire & Mehl (2008), who predicts a person's month best: the person, a close other, or the Echo? Design. The same envelope claims are also rated in advance by the traveler (sealed) and a witness; outcomes judged by the archive.

Study 6 — What does telling do to the teller?

Question. Does daily life-story work change self-concept clarity or wellbeing, for better or worse?Design. Validated questionnaires at the start, month 6 and month 12, by people trained to interpret them, with a waitlist comparison. Why it matters. If the instrument changes the person, the model is chasing a moving target, and we owe people that knowledge.

Study 7 — Does revisiting a memory make it bedrock?

Question. Does a detail revisited at spaced intervals over three months come back more consistently than one never revisited, and does the Echo that leans on bedrock score higher? Design. Within-person: memories randomly assigned to a revisit schedule or to none; consistency at the end of the schedule as the outcome; fidelity trials restricted to bedrock versus the rest. Prediction. The spacing effect holds for autobiographical detail, and fidelity on bedrock exceeds fidelity on drift.

Study 8 — Does your own place beat a neutral one at bedtime?

Question. For one person, do nights ending in a scene from their own archive go better than nights ending in a neutral scene? Design. An n-of-1 trial (Lillie et al., 2011): nights alternated at random between the two, morning rating as the outcome, an optional objective outcome from a wearable where the person allows it (de Zambotti et al., 2019). Prediction. Own-scene nights rate higher for most people, and the size of the difference varies more between people than within them.

11. What would falsify our claims

A protocol that cannot fail is not a protocol. Ours would be in trouble if any of the following held, and we would publish it.

  • Witness tests settling well above 50% after a full Epic — the Echo would simply not resemble anyone.
  • Passive fidelity failing to improve as the archive grows — the material would not be doing the work.
  • Dream closeness scores clustering low after correction — the reconstruction would be a literary exercise.
  • Inferred memories being corrected or deleted by travelers at a high rate — we would be inventing people.
  • Envelope accuracy failing to beat kind-level base rates, or a Brier score that does not fall below 0.25 over time — the Echo would be guessing.
  • Witnesses telling the Echo apart on new sparks while failing to on archive questions — the kernel would not generalise, and the Echo would only be reciting.
  • Low agreement (κ near zero) between travelers and the archive on envelope lines — one of the two verdicts would be meaningless.

12. What we do not claim

  • No consciousness is copied, uploaded or preserved. Nobody knows how to do that.
  • No study has been run on Selione itself. If one is, the protocol will be published before the results.
  • None of the researchers cited has reviewed, endorsed or been involved in this work.
  • Fidelity measures resemblance on what was tested. It is not a measure of identity, and never of a soul.
  • Selione is not a treatment and is not a substitute for care or for grief support.

13. Open questions

  • Does a year of envelopes make an Echo more accurate about its traveler, or does the traveler change faster than the Echo learns?
  • What does long-term use do to a bereaved person? This needs bereavement specialists, not us, and it gates everything that opens after a death.
  • How much of what a person tells us is already a performance for the machine, and does that change over months?
  • Should an Echo update its opinions after its person's death, and if so, within what limits?
  • What is the right unit of resemblance for a person who contradicts themselves — the average, or the contradiction?

14. Data, consent and partnerships

Research participation is a separate, revocable yes, off by default. It adds a traveler's numbers — scores, categories, counts, delays, stages; never words, never a line of the register, never a memory — to aggregates published only for groups of at least twenty people. Withdrawal removes the person from every aggregate computed afterwards. Any model built from those aggregates serves prediction, never targeting, and a scientific partner reviews a publication before it appears. We never sell or license data. What is published lives at Research · data. Researchers and funders who want to work on these questions can reach us through Research · partners.

References

  1. Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337–351.
  2. Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.
  3. Butler, R. N. (1963). The life review: an interpretation of reminiscence in the aged. Psychiatry, 26, 65–76.
  4. Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.
  5. de Zambotti, M., Cellini, N., Goldstone, A., Colrain, I. M., & Baker, F. C. (2019). Wearable sleep technology in clinical and research settings. Medicine & Science in Sports & Exercise, 51(7), 1538–1557.
  6. Donnellan, M. B., Oswald, F. L., Baird, B. M., & Lucas, R. E. (2006). The Mini-IPIP scales: Tiny-yet-effective measures of the Big Five factors of personality. Psychological Assessment, 18(2), 192–203.
  7. Lillie, E. O., Patay, B., Diamant, J., Issell, B., Topol, E. J., & Schork, N. J. (2011). The n-of-1 clinical trial: the ultimate strategy for individualizing medicine? Personalized Medicine, 8(2), 161–173.
  8. Lindeman, M., & Verkasalo, M. (2005). Measuring values with the Short Schwartz's Value Survey. Journal of Personality Assessment, 85(2), 170–178.
  9. Nielsen, T. A., Kuiken, D., Alain, G., Stenstrom, P., & Powell, R. A. (2004). Immediate and delayed incorporations of events into dreams. Journal of Sleep Research, 13(4), 327–336.
  10. Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.
  11. Clark, A. (2013). Whatever next? Predictive brains, situated agents, and the future of cognitive science. Behavioral and Brain Sciences, 36(3), 181–204.
  12. Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46.
  13. Conway, M. A., & Pleydell-Pearce, C. W. (2000). The construction of autobiographical memories in the self-memory system. Psychological Review, 107(2), 261–288.
  14. Diekelmann, S., & Born, J. (2010). The memory function of sleep. Nature Reviews Neuroscience, 11(2), 114–126.
  15. Ericsson, K. A., & Simon, H. A. (1980). Verbal reports as data. Psychological Review, 87(3), 215–251.
  16. Gneiting, T., & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378.
  17. Haber, S., & Stornetta, W. S. (1991). How to time-stamp a digital document. Journal of Cryptology, 3(2), 99–111.
  18. Kahneman, D., & Klein, G. (2009). Conditions for intuitive expertise: A failure to disagree. American Psychologist, 64(6), 515–526.
  19. Klass, D., Silverman, P. R., & Nickman, S. L. (Eds.) (1996). Continuing Bonds: New Understandings of Grief. Taylor & Francis.
  20. Loftus, E. F. (2005). Planting misinformation in the human mind: A 30-year investigation of the malleability of memory. Learning & Memory, 12(4), 361–366.
  21. Mellers, B., et al. (2014). Psychological strategies for winning a geopolitical forecasting tournament. Psychological Science, 25(5), 1106–1115.
  22. Merton, R. K. (1948). The self-fulfilling prophecy. The Antioch Review, 8(2), 193–210.
  23. Murphy, A. H. (1973). A new vector partition of the probability score. Journal of Applied Meteorology, 12(4), 595–600.
  24. Nisbett, R. E., & Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231–259.
  25. Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.
  26. Park, G., Schwartz, H. A., Eichstaedt, J. C., et al. (2015). Automatic personality assessment through social media language. Journal of Personality and Social Psychology, 108(6), 934–952.
  27. Park, J. S., Zou, C. Q., Shaw, A., et al. (2024). Generative agent simulations of 1,000 people. arXiv:2411.10109.
  28. Polanyi, M. (1966). The Tacit Dimension. Doubleday.
  29. Prigerson, H. G., Horowitz, M. J., Jacobs, S. C., et al. (2009). Prolonged grief disorder: Psychometric validation of criteria proposed for DSM-V and ICD-11. PLoS Medicine, 6(8), e1000121.
  30. Schacter, D. L. (1999). The seven sins of memory: Insights from psychology and cognitive neuroscience. American Psychologist, 54(3), 182–203.
  31. Stroebe, M., & Schut, H. (1999). The dual process model of coping with bereavement: Rationale and description. Death Studies, 23(3), 197–224.
  32. Vazire, S., & Mehl, M. R. (2008). Knowing me, knowing you: The accuracy and unique predictive validity of self-ratings and other-ratings of daily behavior. Journal of Personality and Social Psychology, 95(5), 1202–1216.
  33. Westerhof, G. J., & Bohlmeijer, E. T. (2014). Celebrating fifty years of research and applications in reminiscence and life review: State of the art and new directions. Journal of Aging Studies, 29, 107–114.
  34. Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association, 22(158), 209–212.
  35. Youyou, W., Kosinski, M., & Stillwell, D. (2015). Computer-based personality judgments are more accurate than those made by humans. Proceedings of the National Academy of Sciences, 112(4), 1036–1040.

If you work in autobiographical memory, forecasting, bereavement, or the ethics of digital afterlives and want to argue with any of this, we would rather hear it now than after launch: write to us.