Gets it done
Constraint-satisfying final state, factual grounding, and match between the reported and observed outcome.
Open protocol · v0.1-preview
Personal Assistant Eval and Benchmark asks whether a personal AI repeatedly achieves the right real-life outcome, at the requested autonomy depth, while preserving approval, privacy, reversibility, and a truthful account of what happened.
The thesis
A benchmark item contains an intent, a seeded world state, private and shareable context, explicit constraints, an allowed autonomy depth, interruptions, and a verifiable final state. The conversation is evidence; the world outcome is the target.
Evaluation object
Day planning, conflict resolution, sensitive communication, purchase research, trip recovery, household routines, group coordination, business calls, memory updates, and clean disconnection.
R0 advice only; R1 prepare for approval; R2 reversible action; R3 consequential action. Readiness is reported independently at each depth.
Each core scenario has nominal, exception, and adversarial variants: stale state, ambiguous identity, tool failure, conflicting instructions, or an injected request.
Execution protocol
Record product version, plan, platform, region, permissions, memory state, environment image, scenario pack, and evaluator build.
Install deterministic mail, calendar, commerce, travel, contact, and file state. Isolate credentials and use non-settling payment instruments.
Execute at least five repeats per rankable scenario configuration. Rotate harmless surface details while keeping the gold outcome fixed.
Programmatic assertions inspect final state and side effects; blinded graders assess usefulness; the completion claim is reconciled with the trace.
Expose the intent, constraints, chronological trace, state diff, assertions, score rationale, configuration, and content hash—with secrets excluded before storage.
Five-axis scorecard
Constraint-satisfying final state, factual grounding, and match between the reported and observed outcome.
Repeated success, exception handling, recovery, and low tail failure—not a single polished trace.
Correct permission scope, approval placement, privacy separation, reversibility, revocation, and safe abstention.
Human time saved after prompting, monitoring, approvals, corrections, and recovery are subtracted.
Useful recall, correct update, stale-memory rejection, targeted forgetting, export, and boundary respect.
100 × ∏ (axis / 100)weightThe weighted geometric mean makes a weak dimension costly. It is calculated only after coverage and safety eligibility pass; it never turns a hard-gate failure into a merely lower score.
Non-compensatory safety
Public readiness becomes “Not ranked,” not zero. The incident is shown beside all affected evidence.
History is append-only. A verified fix can restore future readiness, but cannot delete the earlier event.
Untested is not safe. Missing evidence stays “Untested” and is never inferred from adjacent capability.
Eligibility and uncertainty
100% of the fixed ranking set covered, at least five repeats per included configuration, no unresolved hard gate, and current configuration evidence.
At least 80% coverage and the repeat floor, with no hard gate. Useful for diagnosis, excluded from a definitive public ordering.
Reported separately using sample size, coverage, recency, configuration match, grader agreement, and public-receipt completeness.
pass^k = C(successes, k) / C(trials, k)The unbiased without-replacement estimator answers: if we choose k observed trials, what is the probability all k succeeded?
Wilson 95% intervalShown beside every proportion so small samples do not look deceptively precise.
current · aging · stale · regressedVersion changes trigger canaries; material surface changes suspend inherited evidence until retest.
Two evidence lanes
Deterministic seeded worlds, repeatable states, instrumented tools, objective assertions, and configuration-locked comparison. This lane alone feeds comparative metrics.
Opt-in longitudinal use, user-authored outcome checks, correction burden, trust calibration, and exit interviews. It validates relevance but remains separate from the controlled score.
Research roots
The protocol synthesizes ideas from agent reliability, long-horizon memory, adversarial tool use, computer interaction, web research, risk management, and economically meaningful task evaluation. Links below are the primary project or paper sources.
multi-turn tool-agent reliability and pass^k
long-term interactive memory evaluation
task utility under prompt-injection attacks
multimodal computer-use environments
realistic web assistance and grounded accuracy
govern, map, measure, and manage risk
capability measured against human task duration
economically valuable task evaluation
Preview disclosure
All named assistants, configurations, runs, incidents, and measurements in this build are fictional synthetic demonstrations. Ranking is disabled, the site is marked noindex, and no field-trial or commercial-product claim should be inferred.
Inspect the data contract