If Layer 1 validates — if passive longitudinal observation produces calibrated, target-specific lift beyond the R1 population-prior and R2 own-routine baselines — the next question is how to observe a person continuously over months to years. That is a different optimization problem from any current product. Every question below is answerable with the apparatus already specified: fix the task set, vary the sensing arm, read the evidence-efficiency curve.
Manuscript and benchmark release V1.0 — pre-pilot protocol release. Synthetic harness only. No human-subject results.
Each rung presupposes the ones below it, and the program can stall or fail at any rung — a failure that is itself an informative result. Nothing above mechanics is claimed today.
The staged roadmap. Phase 0 — synthetic harness (complete): proves the pipeline, nothing about people. Phase 1 — five-person, thirty-day feasibility pilot (next): proves the protocol runs on real, consented lives and yields variance estimates; not powered for effects. Phase 2 — powered target-specificity study, sized by simulation from Phase-1 estimates. Phase 3 — evidence-tier and model–evidence frontier study. Phase 4 — cross-task transfer. Phase 5 — downstream assistance utility. No phase borrows the license of a later one, and sample sizes for Phases 2–5 are outputs of Phase 1, not promises.
Status. These are open research questions, stated as formulation. None has been answered; no experiments have run. The proposed pre-pilot feasibility study (five consenting participants, thirty days, audio-first, daily sealed predictions) is designed to validate only that the harness can run — it is not powered to answer any question on this page.
The synthetic reference path exercises the full loop — sealing, resolution, scoring, permutation. Run it.
Fix the task set, vary the sensing arm (tier, schedule, device), and report Skill vs R1/R2 per arm. The playbook.
Reality Mining, digital phenotyping, GLOBEM, PULSE, Ego4D, and the memory-benchmark family each solved part of this problem. The reviewed landscape maps what each measures and how it could participate.
Evaluate a capture device against the hardware framework: coverage, evidence, research fitness — and pre-register an evidence-efficiency estimate.