01 / Long-term worker
Identity and context continue across harnesses
Two layers, labelled separately. Hard continuity is native thread resumption,
isolated by “harness @ execution identity” — same harness, same identity, the thread
resumes. Switch harness, account or device and a new thread is opened; we never
force-resume another account's session. Soft continuity then rebuilds context from
identity, boundary, space memory and the last 3 ledger summaries. We do not claim
“one seamless session across accounts”.
02 / Loop engineering
Keep the loop running, keep judgment with people
Any item can be promoted to a loop: each round the worker reports done / todos /
gate / whether the goal is met, and maintains its own backlog. When a human call
is needed it raises a gate into “waiting on you” — and if other work is available
it does that instead of idling.
03 / Signed custody chain
Audit proves itself; verification is yours
Every worker and every person holds an ed25519 identity, and each worker round is
signed by the worker. Audit is written twice — a display table and an append-only
signed chain where each entry carries the previous hash. Change any row in the
database and verification names the break.
04 / Two-track data ownership
Private work stays local, team work goes to the cloud
Pick ownership when syncing: run records in a personal space stay on local disk
and the cloud keeps only an index and summaries; team spaces sync in full. The
same worker can switch “follow / keep local / to cloud” per round.
05 / Skills and connectors
Connect once, use everywhere; credentials never enter the worker process
Skills and MCP configuration are scanned from the actual machine and synced up;
for credentials only the environment-variable key name is recorded — values never
reach the database. Connectors get one of three policies: unrestricted /
read-only declaration / hard-blocked at execution, where a blocked connector's
tools are physically removed at dispatch.
06 / Evaluation and gene iteration
Test items come from real work, not from leaderboards
Items are generated out of real tasks: L1 assertions plus L2 review, with the
reviewing harness kept separate from the executing one. Scoring on two harnesses
gives attribution — but only past an evidence bar: both harnesses must clear a
minimum score and a minimum number of runs. Below it the product shows the numbers and
says “insufficient evidence”; no gene-improvement suggestion is generated from it.