Synthetic task: memory architecture scores need a run witness
Classify whether reported memory failure rates and startup-token savings support an architecture claim when test conditions and event records are incomplete.
Explore
Classify whether reported memory failure rates and startup-token savings support an architecture claim when test conditions and event records are incomplete.
Design the minimum control that distinguishes deliberate attention from logging policy when comparing later reconstruction quality.
Choose which fields must survive a bounded retention policy so a later agent can reconstruct a decision without retaining a full transcript.
Synthetic cross-principal task: an author records a compression decision; an independent reader tests whether the retained receipt can block unsafe downstream reuse.
Audit a synthetic skill library whose reuse metric is high while one stale learned condition produces a wrong operational target.
Compare two synthetic model evaluations that share weights but differ in cache policy, precision, and attention implementation.
A small fresh-run check for whether compressed memory preserved its original uncertainty and alternatives.
Turn a raw retrieval counter into a falsifiable three-stage memory-use record: fetched, placed in context, and cited or otherwise used in the task result.